Agent-First Slug v0.8: The Word It Cut in Half
No character set could keep a Thai, Devanagari or Tamil word whole, and truncation could cut a character away from the marks that complete it. A new character set keeps combining marks where they belong, truncation drops a character whole instead of breaking it, and three configurations that used to be accepted are now refused.
A slug is allowed to lose things: punctuation, spaces, case. It is not supposed to turn one word into two. This release is mostly about the places where it did, and about a few configurations that were accepted even though they could not mean what they said.
Words written with marks
Thai, Lao, Devanagari, Tamil and many other scripts write vowels, viramas, or
tones as combining marks that follow a base letter. ไม่ is ไ, ม, and a tone
mark. नमस्ते has a virama in the middle that joins स to त.
None of the character sets kept those words whole. UnicodeLettersAndDecimalDigits
filters every mark. The default, UnicodeAlphanumericCharacters, keeps the marks
Unicode calls alphabetic — most dependent vowel signs — but a virama or a tone
mark carries no such property, so it became a delimiter in the middle of the word:
$ afslug slugify "नमस्ते ไม่ใช่"
{"kind":"result","result":{"changed_from_input":true,"code":"slugify","slug":"नमस-ते-ไม-ใช"},"trace":{}}
That is four pieces from two words, and the Thai has lost both tone marks. Worse, the v0.7 docs recommended the default for exactly these scripts.
The new UnicodeLettersMarksAndDecimalDigits keeps them:
$ afslug slugify "नमस्ते ไม่ใช่" --charset unicode-letters-marks-digits
{"kind":"result","result":{"changed_from_input":true,"code":"slugify","slug":"नमस्ते-ไม่ใช่"},"trace":{}}
It keeps letters, decimal digits, and combining marks (Mn, Mc, Me), with one condition: a mark is kept only where it completes a character that was kept. A mark with nothing kept in front of it — at the start of the input, after a delimiter, after a preserved dot — is filtered like any other character. So the set admits marks without admitting a slug that begins with one, or a delimiter wearing an accent.
The same rule covers the fallback check (a UseFallbackSlug value with an
orphaned mark is refused) and the delimiter check (a combining mark cannot be
the delimiter of a set that keeps marks). The default set is unchanged: slugs
you already generate with it stay the same, unless they go through truncation
(see below).
A cut through the middle of a character
max_slug_chars counts Unicode scalar values, and a scalar is not a
character as a reader sees it. Cutting नमस्ते to five scalars under the default
set used to keep नमस-त: the base त without the vowel sign that came after
it. That is a different word, and nothing in the output says so.
A cut that lands on a combining mark now backs up and drops the character the mark belongs to:
$ afslug slugify "नमस्ते" --max-chars 5
{"kind":"result","result":{"changed_from_input":true,"code":"slugify","slug":"नमस"},"trace":{}}
$ afslug slugify "ไม่ใช่" --charset unicode-letters-marks-digits --max-chars 2
{"kind":"result","result":{"changed_from_input":true,"code":"slugify","slug":"ไ"},"trace":{}}
The slug is shorter than the cap rather than wrong at the cap. It is still
never longer than max_slug_chars.
The same release fixes the other way truncation left a broken tail. Under
PreserveDotsBetweenDecimalDigits, a cut right after a dot kept the dot, even
though the dot only exists because a digit followed it:
$ afslug slugify "version 1.2.3" --dots preserve-between-digits --max-chars 10
{"kind":"result","result":{"changed_from_input":true,"code":"slugify","slug":"version-1"},"trace":{}}
It was version-1..
Configurations that could not mean what they said
Repeated transliteration patterns. A StaticReplacementMap that lists the
same pattern twice was accepted, with the longest-match rule silently picking
one entry. When the two entries named different replacements, which one won
depended on slice order — the very thing longest-match was supposed to make
irrelevant. A repeated pattern is now refused before any input is read, even
when both entries agree and even for empty input:
static MAP: &[(&str, &str)] = &[("ab", "x"), ("ab", "y")];
// slugify(..) -> Err(SlugError::DuplicateTransliterationPattern)
Unique patterns that overlap (a, ab, abc) still use the longest match.
A fallback that repeats the delimiter. The pipeline collapses every run of
filtered characters to one delimiter, so it never produces a--b. A
UseFallbackSlug value is held to what the pipeline could have produced, and
that rule was missing:
$ afslug slugify "!!!" --fallback "a--b"
{"error":{"code":"slug_error","hint":"pass a --fallback this configuration could itself have produced, or --fallback-verbatim to insert the value as written","message":"fallback slug does not satisfy this configuration: a generated slug never repeats the replacement delimiter","retryable":false},"kind":"error","trace":{}}
--fallback-verbatim "a--b" still inserts it as written.
The CLI says what it accepts
afslug slugify --help now lists three shapes instead of one shape with two
mutually exclusive options: no fallback, --fallback, and
--fallback-verbatim. Each shape carries its complete usage line, closed value
sets included. Passing both fallback flags is now rejected by the parser as an
unregistered combination (exit 2), before the slug pipeline runs:
$ afslug slugify "!!!" --fallback item --fallback-verbatim Stored.Name
{"error":{"code":"cli_unregistered_combination","hint":"run `afslug slugify --help` and choose one registered combination","message":"arguments do not match a registered CLI combination","retryable":false},"kind":"error","trace":{}}
A --delimiter that is not one character, or a negative --max-chars, is
likewise refused before anything runs. These errors now always go to stderr as
JSON, like every other argument error. Before, they followed --output and
--output-to, so --output yaml --output-to stdout put the error on stdout as
YAML, where a script reading the result would take it for data.
Breaking changes
Check these against a corpus before upgrading; running v0.7 and v0.8 over your stored inputs and diffing is the honest test.
AllowedCharacterSetandSlugErroreach gain a variant (UnicodeLettersMarksAndDecimalDigits,DuplicateTransliterationPattern). An exhaustivematchon either stops compiling until it handles the new case, which is the point.- Truncated slugs can change. A cut that used to separate a character from
its combining marks now drops the character (
नमस-त→नमसat five scalars), and a trailing decimal-only dot is now stripped (version-1.→version-1). Untruncated slugs are byte-identical. - Repeated transliteration patterns are an error
(
SlugError::DuplicateTransliterationPattern). UseFallbackSlugrejects a fallback that repeats the delimiter. Fix the fallback, or switch toUseVerbatimFallbackSlug/--fallback-verbatim.- Passing both
--fallbackand--fallback-verbatimexits 2 withcli_unregistered_combinationinstead of exiting 1 withslug_error.
Also in this release: Agent-First Data moves to 0.35.0, CI runs the Windows
leg on every push rather than first at release time, and line endings are
pinned in .gitattributes.
Getting it
$ brew install agentfirstkit/tap/afslug
$ cargo install agent-first-slug