Specification
Self-describing structured data for AI agents and humans.
Field names encode units and semantics. Agents read latency_ms and know milliseconds, api_key_secret and know to redact — no external schema needed.
Overview
Agent-First Data has three parts:
- Naming Convention (required) — encode units and semantics in field names
- Output Processing (required) — suffix-driven formatting and automatic secret protection
- Protocol Template (optional) — structured format with
kindand its matching payload (required) andtrace(recommended)
Parts 1 and 2 are the core. Part 3 is optional — a recommended structure that works well with Parts 1 and 2, but you can use AFDATA naming with any JSON structure (REST APIs, GraphQL, databases, etc.).
Four SDKs, one contract. Naming, output processing, the protocol template,
and the established core helpers are implemented in Rust, Go, Python, and
TypeScript. Rust is the reference compiler for the new closed-world
cli-spec-v1 registry and cli-help-v2; the other SDK compilers follow the
same serialized model and fixtures. The Rust crate also bundles
skill/skill-admin, stream-redirect, and tracing integration. The exact
shared legacy surface is enumerated in spec/api-surface.json.
Jump to:
Quick Reference: All Suffixes
| Category | Suffixes | Plain example |
|---|---|---|
| Duration | _ns, _us, _ms, _s, _minutes, _hours, _days | latency_ms: 1280 → latency: 1.28s |
| Timestamps | _epoch_ns, _epoch_ms, _epoch_s, _rfc3339 | created_at_epoch_ms: 1707868800000 → created_at: 2024-02-14T... |
| Size | _bytes | file_size_bytes: 5242880 → file_size: 5.0MiB |
| Currency | _msats, _sats, _usd_cents, _eur_cents, _jpy, _{code}_cents, _{code}_micro | price_usd_cents: 999 → price: $9.99 |
| String formats | _bcp47, _utc_offset, _rfc3339_date, _rfc3339_time | language_bcp47: "zh-CN", invoice_due_rfc3339_date: "2026-06-13" |
| Other | _percent, _secret, _url | cpu_percent: 85 → cpu: 85% |
JSON and YAML are structure-preserving: both keep original keys, scalar types, and numeric semantics after redaction — YAML is just JSON’s data rendered as a multi-line document. In default Plain only: formatting suffixes are stripped from keys (value already encodes the unit) and values are formatted for readability. (_url, _bcp47, _utc_offset, _rfc3339_date, and _rfc3339_time are never stripped, even in Plain.)
Secret protection: All three formats automatically redact _secret fields and scrub secret components (userinfo password, secret-named query params) inside _url field values.
Boundary: AFDATA names communicate local field semantics. They do not replace schemas for required fields, enum values, numeric ranges, object shapes, or cross-field validation. Use JSON Schema, OpenAPI, database constraints, or typed APIs for those guarantees.
Part 1: Naming Convention
Applies to all structured data: JSON, YAML, TOML, CLI arguments, environment variables, config files, database columns, HTTP payload fields, log fields.
Design rules
- Name conveys meaning. A reader should understand the field’s purpose from the name alone, without seeing surrounding context or documentation.
datacould be anything —request_body,search_results,cached_responsesay exactly what it contains. - Unit in suffix. If a numeric value has a unit, encode the unit in the field name suffix.
- Secrets marked. If a value is sensitive, end the field name with
_secret. - Obvious needs no suffix. If the meaning is obvious from the name alone, no suffix is needed.
- Self-contained. Never rely on external metadata, companion fields, or documentation to convey what a field contains.
Suffixes
afdata lint type-checks every suffix below, but only for a present, non-null value: a null value means the field is absent/unset, not present-with-the-wrong-type, so it is exempt from the suffix’s type constraint. Absence may be written as an omitted key or as an explicit null — both are valid; a bare law_repeal_at_epoch_s: null is not a suffix_type_mismatch.
Schema input. An object’s properties map is read as a JSON Schema property map, not as data: each key is checked as a field name and its value as that field’s schema, so duration_ms: {"type": "integer"} satisfies the _ms check instead of being read as an object-valued duration. What is checked there is the declared type — it must admit the suffix’s type, through type arrays ({"type": ["integer", "null"]} passes), anyOf/oneOf (any branch), and allOf (every branch), while {"type": "string"} under duration_ms is a suffix_type_mismatch. A schema that constrains no primitive type at that location — a $ref, a bare const, a boolean schema — is left alone rather than guessed at, and the value-level checks below (RFC 3339 structure, BCP 47 well-formedness, URL shape, non-negativity) have no schema counterpart: a _rfc3339 property is only required to declare a string. Because a JSON Schema is always an object or a boolean, a properties entry whose value is a scalar or an array is data rather than a schema and keeps the ordinary runtime check — {"properties": {"timeout_ms": "5000"}} is still a suffix_type_mismatch. A _secret property’s default, example, and examples must each be absent (null) or already redacted (***); a literal there is a secret_schema_value_exposed error, since a schema is published far more widely than the value it describes.
Duration
| Suffix | Unit | Example |
|---|---|---|
_ns | nanoseconds | gc_pause_ns: 450000 |
_us | microseconds | query_us: 830 |
_ms | milliseconds | latency_ms: 142 |
_s | seconds | dns_ttl_s: 3600 |
_minutes | minutes | session_timeout_minutes: 30 |
_hours | hours | token_validity_hours: 24 |
_days | days | cert_validity_days: 365 |
Timestamps
| Suffix | Format | Example |
|---|---|---|
_epoch_ns | nanoseconds since Unix epoch as a decimal string | created_epoch_ns: "1707868800000000000" |
_epoch_ms | milliseconds since Unix epoch | created_at_epoch_ms: 1707868800000 |
_epoch_s | seconds since Unix epoch | cached_epoch_s: 1707868800 |
_rfc3339 | RFC 3339 date-time string | expires_rfc3339: "2026-02-14T10:30:00Z" |
_epoch_s and _epoch_ms use JSON integers. Current-era _epoch_ns values exceed JSON’s cross-language safe integer range, so _epoch_ns uses a decimal string.
*_rfc3339 names an RFC 3339 date-time: a full-date, a T (or t) separator, a partial-time with optional fractional seconds, and a mandatory time-offset — either Z/z or ±HH:MM. AFDATA validates this structure: 2026-02-14T10:30:00Z and 2026-02-14T10:30:00.5+08:00 pass, while a bare 2026-02-14T10:30:00 (no offset), a space separator, or a trailing IANA name is rejected. A leap second (:60) is rejected, as for _rfc3339_time. For an instant with no wall-clock form, prefer _epoch_*. The is_valid_rfc3339 helper and afdata lint apply exactly this check.
Strict string formats
These suffixes identify strings with a strict external format. They are semantic field-name conventions, not Plain formatting suffixes: Plain’s default readable output keeps the full key and raw string value for these (YAML always keeps every key and value as-is, for every suffix).
| Suffix | Format | Example |
|---|---|---|
_bcp47 | BCP-47 language tag string | language_bcp47: "zh-CN" |
_utc_offset | fixed UTC offset string | timezone_utc_offset: "+08:00" |
_rfc3339_date | RFC 3339 full-date string | invoice_due_rfc3339_date: "2026-06-13" |
_rfc3339_time | RFC 3339 partial-time string | market_open_rfc3339_time: "09:30:00" |
*_bcp47 names a field whose string value is a BCP-47 language tag, such as language_bcp47: "zh-CN" or content_language_bcp47: "en-US". AFDATA validates the tag’s structure — hyphen-separated ASCII-alphanumeric subtags with a 2–3 letter primary language subtag — which rejects the common POSIX-locale form zh_CN (underscore), an over-long primary like chinese, and other malformed tags. It does not check the IANA subtag registry, so a structurally well-formed but unregistered tag like zz-ZZ still passes; a tool needing that stronger guarantee validates further. The is_valid_bcp47 helper and afdata lint apply exactly this structural check.
*_utc_offset names a fixed offset from UTC. Canonical persisted and structured output values are "UTC" or ±HH:MM, with HH in 00..23 and MM in 00..59; zero offsets normalize to "UTC". Examples: timezone_utc_offset: "+08:00", report_utc_offset: "-05:00". This is intentionally not an IANA timezone name: do not use Asia/Shanghai, America/Los_Angeles, DST rules, or timezone databases in this field.
*_rfc3339_date names an RFC 3339 full-date string: exactly YYYY-MM-DD, such as invoice_due_rfc3339_date: "2026-06-13". It is a calendar date, not an instant, and it does not imply any time, offset, or timezone.
*_rfc3339_time names an RFC 3339 partial-time string: exactly HH:MM:SS with optional fractional seconds, such as market_open_rfc3339_time: "09:30:00" or "09:30:00.123". It is a time-of-day, not an instant. It MUST NOT include Z, ±HH:MM, an IANA timezone, or any other timezone annotation; a time without a date cannot be resolved through timezone/DST rules. Use _rfc3339 or _epoch_* for instants.
AFDATA core does not define a companion timezone-name field. If a future tool needs to preserve IANA timezone semantics with a timestamp, prefer a self-contained standard value such as RFC 9557 rather than pairing a date/time field with a separate timezone-name field.
Tools should avoid magic string sentinels such as "auto" inside strict-format fields. If a tool needs auto/default behavior, define that in the tool’s own config semantics rather than as an AFDATA-wide rule.
Size
| Suffix | Value type | Usage | Example |
|---|---|---|---|
_bytes | non-negative integer | Output, APIs, config | payload_bytes: 456789 |
Byte sizes are always integer _bytes, in inputs and outputs alike. AFDATA has no unit-in-value size string: a field like buffer_size: "10MiB" would force every reader to parse units before comparing or summing, and the _size name collides with count-style fields (page_size, batch_size, pool_size) that are quantities, not byte sizes. Pick the unit once at schema-design time and encode it in the key, exactly as durations do (timeout_s, not timeout: "30 minutes").
In Plain output, _bytes values auto-scale to human-readable format (5.0MiB, 2.0GiB). JSON and YAML keep the raw integer.
Percentage
| Suffix | Unit | Example |
|---|---|---|
_percent | percentage | cpu_percent: 85 |
The value is in units of percent: 1 means 1%, 0.2 means 0.2%, 85 means 85%. The % unit is exactly 0.01, so a _percent value is never a 0–1 fraction — a producer holding a ratio such as 0.999 multiplies by 100 (99.9) before writing it. Decimals, negatives, and values above 100 are all valid (a multi-core CPU can report 800, a delta can be -5); AFDATA fixes no global range. Writing 0.85 when you mean 85% is a unit-conversion error on the producer’s side, not an ambiguity in the field convention.
Currency
Bitcoin:
| Suffix | Unit | Example |
|---|---|---|
_msats | millisatoshis as an integer or decimal integer string | balance_msats: 97900 |
_sats | satoshis as an integer or decimal integer string | withdrawn_sats: 1234 |
AFDATA does not define a floating _btc suffix. Use integer _sats or _msats instead.
Fiat — _{iso4217}_cents for currencies with 1/100 subdivision, _{iso4217} for currencies without (JPY). Always signed integers: positive and negative values represent amounts in the same unit, so refunds, credits, reversals, and deltas do not need a second suffix:
| Suffix | Unit | Example |
|---|---|---|
_usd_cents | US dollar cents | price_usd_cents: 999 |
_eur_cents | euro cents | price_eur_cents: 850 |
_thb_cents | Thai baht 1/100 | fare_thb_cents: 15050 |
_jpy | Japanese yen (no minor unit) | price_jpy: 1500 |
Stablecoins follow the same _{code}_cents pattern: deposit_usdt_cents: 1000, payout_usdc_cents: 500.
Sub-cent precision — _{code}_micro for integer micro-units, one millionth (10⁻⁶) of the major unit:
| Suffix | Unit | Example |
|---|---|---|
_{code}_micro | millionths of one major unit | cost_usd_micro: 170000 (= $0.17) |
_{code}_micro is the fiat analog of _msats: when cents are too coarse (per-token LLM pricing, metered API costs, unit-economics accounting), do not switch to decimal cents — move to a smaller integer unit. Values are always signed integers. Use _{code}_cents for user-facing amounts and _{code}_micro for high-precision internal accounting.
Sensitive
| Suffix | Handling | Example |
|---|---|---|
_secret | redact the entire value/subtree to ***, except null, which stays null | api_key_secret: "sk-or-v1-abc..." |
_url | redact secret components inside the URL value (userinfo password, secret-named query params); the rest of the URL is preserved | callback_url: "https://h/cb?code_secret=..." |
All CLI output formats (JSON, YAML, Plain) automatically redact _secret fields. Any _secret value — scalar, object, or array — becomes the scalar string ***, so a secret-marked container never leaks through JSON, YAML, Plain, or collision fallback. The one exception is null, which is preserved: a null secret is an absent secret, and masking it would manufacture the appearance of a configured credential — a reader cannot tell ***-because-set from ***-because-null, so a tool showing its own config would report every unset secret as configured. Redaction hides a value that exists; it does not invent one. (An empty string is a value and is still redacted; only null is absence.) Matching recognizes _secret and _SECRET only. Config files always store the real value. For legacy payloads that cannot rename fields to _secret, use OutputOptions.redaction.secret_names (a configured Redactor in Rust/Go, keyword arguments in Python/TS) at serialization time; names match exact field names at any nesting level, with no trim, case folding, hyphen/underscore normalization, globs, regex, or substring matching. Secret-name lists only affect redaction; formatting suffix stripping is a Plain-only concern, controlled by AFDATA suffixes in Plain’s default readable style. AFDATA does not define named redaction profiles; use the default All, secret_names, TraceOnly, or Off deliberately at the serialization boundary. YAML is always schema-preserving, regardless of style. Callers that need schema-preserving Plain rendering can use OutputOptions with PlainStyle::Raw.
Legacy URL fields that cannot be renamed use an explicit url_names list at
the same boundary (Redactor::url_names / Redactor.URLNames,
OutputOptions.url_names, or OutputOptions.redaction.urlNames in native
casing). Matching is the same exact field-name equality as secret_names.
A matching string receives _url treatment; arrays and nested objects recurse
into every string leaf. A secret-named child still redacts its whole value
first. Benign non-URL strings remain unchanged, while malformed
credential-looking leaves retain the normal _url fail-closed behavior.
Guarantee boundary. _secret protects a structured field only when that
value passes through an AFDATA redactor or renderer. It does not scrub process
argv, shell history, /proc, a parent process, free-form prose, or third-party
logs. Keep live secrets out of argv. A CLI that records its invocation must
call redact_argv before logging argv, but that only protects the resulting
structured log; it cannot retroactively protect the process boundary.
The marker *** has exactly one meaning in AFDATA output: a value was redacted because its field name, URL query parameter name, explicit secret_names entry, or exact url_names URL context made it sensitive. It is not used for serialization failures, truncation, unsupported types, or arbitrary “maybe secret” guesses.
Secrets inside URLs
Key-based redaction cannot reach a secret embedded inside a URL string —
token in wss://host/cdp?token=abc is not its own field. Implementations
therefore expose two explicit URL-aware helpers:
redact_url_secrets(url, *, secret_names=())(Python) /redactUrlSecrets(url, options?)(TS) /redact_url_secrets(url)withRedactor{secret_names}.url(url)for custom names (Rust/Go) — returnsurlwith its secret components redacted to***.redact_urls_in_text(text, ...)(Rust/Python) /redactUrlsInText(text, options?)(TS) /RedactURLsInText(text)orRedactor.URLsInText(text)(Go) — recognizes complete scheme-prefixed URL spans inside prose and applies the same URL redactor to those spans only.
The same secret decision as everywhere else applies, to the URL’s query-parameter names: a parameter is redacted iff its (form-decoded) name ends in _secret/_SECRET, or matches an exact entry in secret_names. No built-in list of “sensitive” parameter names exists — a legacy parameter such as ?token= is redacted only when the caller passes secret_names: ["token"], exactly as for legacy field names. Consumers that own the URL should instead rename the parameter to follow the suffix convention (?token_secret=).
⚠️ Common credential-bearing parameters are NOT redacted by default. The userinfo password (
user:pass@host) is always scrubbed structurally, but query parameters are matched by name only. Conventionally-named secret parameters such as?access_token=,?api_key=,?code=,?id_token=,?sig=, or?sessionid=pass through unchanged unless their name ends in_secretor is listed insecret_names. A_urlfield does not make an arbitrary URL safe to log — it scrubs the userinfo password and suffix-named/listed parameters, nothing more. When you own the URL, rename sensitive parameters to the_secretsuffix (?access_token_secret=); when you do not, pass the parameter names viasecret_names.
Independently of the parameter convention, the userinfo password component is always redacted as a structural rule: scheme://user:pass@host → scheme://user:***@host (the username is preserved; a userinfo with no : is left untouched).
Input must be a single URL. The standalone helper processes a string iff it begins with a scheme (^[A-Za-z][A-Za-z0-9+.-]*://) and contains no whitespace; any other string — including a URL embedded in surrounding prose — is returned unchanged. Callers that build messages around a URL redact the URL before interpolating it: format("connect {}: {}", redact_url_secrets(url), err).
Prose scanning is explicit and URL-only. redact_urls_in_text exists for a
message that already contains URLs. It recognizes only structurally complete
scheme URL spans, delegates each span to the standalone URL redactor, and
preserves the surrounding text and punctuation. Structured redaction never
calls it by default, and it never scans arbitrary prose for secret-looking
names or values.
The span grammar. Where a span begins and ends is part of the contract, not an implementation detail: four languages must cut the same bytes out of the same message. Spans are found left to right, and the scan for the next one resumes at the end of the previous one.
- Scheme start. Candidate schemes are looked for at the head of a maximal run
of scheme bytes (
A–Z,a–z,0–9,+,-,.) — a scheme cannot begin in the middle of such a run, so thehttpsinsidefoo-https://h/is not a second candidate and the span starts atf. A run that does not begin with an ASCII letter is retried from its first ASCII letter rather than abandoned:2https://h/?token_secret=xis a URL with a digit glued to its front, and giving up on it fails open, leaving the whole query readable. - Span end. The span runs from the scheme start to the first whitespace
character, the first text delimiter
(
"'<>`,。;!?)》】」』), or the end of the text. Whitespace here is the UnicodeWhite_Spaceproperty, enumerated rather than delegated to each language’s native predicate:U+0009–U+000D,U+0020,U+0085,U+00A0,U+1680,U+2000–U+200A,U+2028,U+2029,U+202F,U+205F,U+3000. The enumeration is the point. Each language’s own predicate covers a different set — JavaScript’s\scountsU+FEFFbut notU+0085, Python’sstr.isspace()countsU+001C–U+001F, Rust and Go count neither — and disagreement is a defect in both directions: ending a span early leaves the rest of the query in the clear, while ending it late pulls the following prose into the URL, where redacting the parameter it landed in deletes it. The set is deliberately generous, because this helper exists to keep a message readable while removing the secret: anything a reader sees as a space ends the URL. Two exclusions are equally deliberate.U+FEFF(zero-width no-break space) is notWhite_Space; treating it as one would cuthttps://h/x<U+FEFF>?token_secret=leakshort and leave the secret in the clear.U+001C–U+001F(the file, group, record, and unit separators) are notWhite_Spaceeither. Both are ordinary URL bytes here, and a span containing them stays one span. Note this is a different question from the single-URL gate above — is this whole string one URL? — which stays ASCII-only; a span the scanner produces contains no whitespace at all, so it satisfies the stricter gate either way. - Trailing trim. Sentence punctuation is then trimmed back off the end of the
span (
.,;!)]},。;!)》】), sosee https://h/?token_secret=abc).keeps its).outside the URL. - Minimum. At least one character must survive after
://once that trim is done. A barehttps://in prose, ors://with nothing attached, is not a span and nothing is replaced.
Each surviving span is handed to the single-URL redactor and every byte outside the spans is copied through unchanged.
Surgical replacement. Only the secret spans (a secret parameter’s value bytes after = up to the next &/#/end; the password bytes after the first : in userinfo up to the authority’s last @) are replaced with the literal ***. Every other byte — scheme, host, path, benign parameters, percent-encoding, ordering — is preserved exactly. Implementations locate the scheme, authority, query, and fragment spans structurally and form-decode parameter names only; they deliberately do not require a URL library to accept the whole input, because malformed-but-credential-bearing URLs must still be scrubbed and cross-language parsers disagree about them. Output equals input outside the redacted spans.
The fragment is parameters too, when it holds them. A fragment written as k=v(&k=v)* is subject to the same secret-parameter rule as the query — the OAuth implicit flow returns tokens there, so a fragment is a routine place for a credential to be, and treating it as opaque text means ?token_secret= is redacted while #token_secret= is not. A fragment whose segments contain no = carries no parameters and is preserved byte for byte (#section, #, #a/b?c). secret_names applies inside the fragment as well.
Automatic application via _url or url_names. Redaction applies URL
treatment to any field whose name ends in _url/_URL, and to an exact legacy
field name explicitly listed in url_names. No unrelated payload string is
scanned: the trigger is still the field name, exactly like _secret. A
configured legacy URL collection recursively applies the treatment to its
string leaves. So final_url, callback_url, and a configured legacy relays
field are scrubbed, while a free-form error or message field is never
touched even if it contains a URL (use the explicit prose helper there).
Off disables this along with all other redaction; TraceOnly scopes it to the
trace subtree. A URL-context value with surrounding whitespace is trimmed
before URL redaction. One that cannot be parsed as a clean scheme-prefixed URL
is replaced with *** rather than silently passing through a likely malformed
secret-bearing value when it carries either internal whitespace or an @
credential sigil — for example a schemeless connection string
user:pass@host:5432/db, which has no scheme anchor for the surgical span
logic. A schemeless, @-free, whitespace-free value (e.g. a relative URL
/cb?page=2) still passes through unchanged. The secret_names list applies
to query-parameter names inside these values as well. (A field carrying both
meanings, e.g. token_url_secret, ends in _secret and so its whole value is
redacted to ***.)
No suffix needed
Fields whose meaning is obvious from the name alone:
- Paths:
redb_path,config_path - Counts:
proof_count,relay_count - Booleans:
search_enabled,forward_pulse - Identifiers:
method,domain,model,backend
(URL-valued fields are the exception: end them in _url so secrets inside the
URL are scrubbed — see the _url suffix above.)
CLI arguments
Same suffixes, kebab-case. An agent reading --help output understands units and sensitivity without documentation:
--timeout-ms 5000 # milliseconds
--cache-ttl-s 3600 # seconds
--max-size-bytes 1048576 # bytes
--api-key-secret SECRET # sensitivity marker only; keep live secrets out of argv
--max-buffer-bytes 1048576 # bytes as an integer, never a "10MiB" string
--port 8080 # no suffix needed — meaning obvious
--verbose # boolean flag — no suffix needed
One spelling per concept. --output is the only format selector AFDATA
recognizes. A --json alias means exactly --output json, so it buys an agent
nothing and costs the application a flag name it may want for something else (a
--json <FILE> input, say) — a pre-parser that claims --json would both
hijack the flag and misread its value as a subcommand. Applications own
--json; AFDATA does not.
Long flags only. Do not define single-letter short flags (-s, -d,
-l). Short flags are ambiguous — -s could be --synapse, --synopsis, or
--source — and agents parsing help output cannot reliably interpret them.
Always use the full --kebab-case form. An application MAY still declare a
short where it carries real meaning: a tool imitating an established interface
may owe its users -h for --host, or -f for --follow on a log command.
Imitation is the bar — a short that merely abbreviates a flag of your own
invention fails it.
Note that a CLI compiled from cli-spec-v1 (Part 3) cannot take that
permission: the registry admits no single-letter aliases or clusters, because
each would add a second spelling to every legal invocation shape. That is a
property of one compiler, not of this naming convention.
The AFDATA help and version handlers claim no shorts of their own: -h/-V would be byte-identical aliases of --help/--version, which buys an agent nothing and misleads the human reaching for them (help answers in the command’s output format, which for an agent-first CLI is JSON, not text). Leaving them unclaimed is also what keeps those letters available to the application. What an application then does with -h is its own decision, not this convention’s.
Kebab → snake mapping. CLI flags map 1:1 to JSON field names by replacing hyphens with underscores. When a CLI tool emits a startup log event (Part 3), the args field uses the snake_case form:
myapp --cache-ttl-s 3600 --api-key-secret sk-xxx --max-size-bytes 1048576
{"kind":"log","log":{"message":"startup","level":"info","event":"startup","args":{"cache_ttl_s":3600,"api_key_secret":"***","max_size_bytes":1048576}},"trace":{}}
---
kind: "log"
log:
args:
api_key_secret: "***"
cache_ttl_s: 3600
max_size_bytes: 1048576
event: "startup"
level: "info"
message: "startup"
trace: {}
The flag name and the JSON/YAML field name tell the same story — the suffix carries the unit even in structure-preserving output, with no separate mapping table or --help prose explaining “timeout is in milliseconds” needed.
Secret flags (--api-key-secret, --database-url-secret) carry the same
naming signal as structured fields, but argv is outside AFDATA’s automatic
redaction boundary. If a tool must record its invocation, it must call
redact_argv before placing argv in a structured startup event or log. That
does not protect shell history, /proc, parent-process inspection, or logs
written by code that bypasses the AFDATA redactor; prefer a non-argv secret
source.
Closed-world invocation registry. An AFDATA CLI MUST be compiled from one
serializable cli-spec-v1 registry. Each command declares local ArgSpecs and
one or more named Combinations. A combination consists only of fixed finite
enum values, explicitly required arguments, optional arguments whose arbitrary
subsets are legal, and an output contract. Every other application argument is
forbidden. An invocation is legal only when it matches exactly one
combination. Zero matches return code:"cli_unregistered_combination",
retryable:false, and exit 2. Multiple matches are a spec build error, not a
runtime priority decision.
The registry is the only truth source for tokenization, type validation,
combination matching, resolved typed values, output planning, and help.
Applications MUST NOT maintain a second clap/argparse/flag definition,
post-parse requires/conflicts policy, ignored compatibility option, or
application global. The full command path comes first, followed by that
command’s options and positionals. Built-in lifecycle/output names are
reserved where AFDATA parses them and nowhere else. --help, --output,
--output-to, --stdout-file, and --stderr-file are parsed at every
command, so no command may declare them. --version and --docs are answered
by the root command alone, so a subcommand MAY declare its own: in
tool release --version 1.2.0 that version is the release’s, and the command
path is fixed before that command’s arguments are parsed, so the two spellings
are never ambiguous. Where a subcommand declares neither, --version and
--docs past the root stay an unregistered combination.
Declared value sources. A string argument MAY declare a non-empty sources
set when its value can be supplied indirectly. The built-in schemes are
env:NAME, file:PATH#DOT_PATH, stdin, fd:N, and prompt;
file+FORMAT:PATH#DOT_PATH names any document format the build supports when
the filename does not. A bare argument remains the literal value, and
literal:VALUE escapes a value beginning with a recognized prefix. Source
arguments cannot declare defaults: a default is a value, never an instruction
to read the host’s environment or filesystem.
The registry classifies sources but does not read them. It MUST reject a
recognized scheme outside the argument’s declared set as
cli_invalid_argument_value before config, filesystem, terminal, network, or
domain I/O. The application reads the classified source only when its command
needs the value. In the Rust reference, ValueSource::read() returns a
String; read_secret() returns a SecretString whose Debug and Display
are *** and whose value is available only through the explicit
expose_secret() boundary. Empty strings are values, not missing sources.
A host MAY add a source it reads itself by declaring a host_scheme. Its name
uses lowercase ASCII letters, digits, and hyphens, begins with a letter, and
must not duplicate another host source or the reserved env, file, stdin,
fd, prompt, or literal names. Its displayed syntax begins with that exact
name: and a value placeholder. A host-only source set is valid. The AFDATA
Rust adapter additionally refuses prompt unless the argument id ends in
_secret, because prompting blocks and suppresses terminal echo.
Help scope and output. Help is generated directly from the same registry
and answers in one round trip: myapp [command] --help returns every
registered shape of that command, each complete, plus the next-level
... --help commands. It does not make the agent solve a requires/conflicts
graph. There is no second help level — the only thing it could omit is the
optional arguments, and a caller that stopped at the first level would neither
have them nor know they existed, which is the defect this design removes.
--docs renders the whole registry as Markdown; there is no recursive help
mode in v2.
JSON/YAML help is a protocol-v1 terminal result with result.code:"help" and a
cli-help-v2 model under result.help; the exact contract is
cli-help-v2.schema.json. Bare help defaults to
JSON. Humans request the equivalent plain catalog with --output plain, which
carries the same shapes, argument meanings, and defaults — a different
rendering, never a smaller one. Each shape’s usage is generated rather than
handwritten and shows fixed, required, optional, and output arguments, naming
a closed value set inline (--output <json|yaml|plain>) and redacting secret
values as placeholders. A fixed argument whose own default already satisfies
it is shown bracketed, because the parser accepts the call without it.
notes and defaults are keyed by the spelling usage uses. Subcommands
sort deterministically; shapes keep their registration order.
Help-v2 is an invocation contract, not a catalog of every domain outcome.
Command-specific idempotency, runtime error codes, side effects, and recovery
semantics stay in the owning tool’s focused documentation and tests; they do
not expand CliSpec or the generic help schema.
Output arguments do not select a business combination and do not participate
in overlap checks. The compiler first selects exactly one application shape,
then validates output only against that combination’s closed raw or
protocol contract. Parse and match failures do not trust unresolved output
arguments: stdout remains empty, stderr receives one strict JSON error event,
and the process exits 2. The failure is named by error.code — one
cli_* code per failure kind, beside the document_* codes and read the same
way — with a safe argument spelling or the failure category in message and
the next command to run in hint. Neither ever carries a raw value. A CLI not
compiled from a registry
reports the generic cli_error; see
protocol-v1.schema.json.
Version output. The registry automatically adds version as a root-only
lifecycle combination. It emits
{"kind":"result","result":{"code":"version","name":"<name>","version":"<semver>"},"trace":{}}
through CliSpec.lifecycle_output. Past the root the spelling is the
application’s: a subcommand that declares --version receives its own
argument, and one that does not rejects the token as
unregistered_combination.
Environment variables
Same suffixes, UPPER_SNAKE_CASE:
DATABASE_URL_SECRET=postgres://user:pass@host/db
CACHE_TTL_S=3600
TOKEN_VALIDITY_HOURS=24
RUST_LOG=info
Config files
Config files follow the same naming suffixes. Agents reading a config file can determine units, formats, and sensitivity without a separate schema.
YAML
openrouter:
api_key_secret: "sk-or-v1-actual-key"
model: "google/gemini-3-flash-preview"
storage:
backend: redb
postgres_url_secret: "postgres://user:pass@host/db"
redb_path: "data.redb"
cache:
dns_ttl_s: 3600
manifest_ttl_s: 300
pricing:
input_msats: 2
output_msats: 12
TOML
[cache]
dns_ttl_s = 3600
manifest_ttl_s = 300
[openrouter]
api_key_secret = "sk-or-v1-actual-key"
model = "google/gemini-3-flash-preview"
Database schemas
Same suffixes in column names. Agents reading a table schema can determine units, formats, and sensitivity without external documentation.
When the database type already carries semantics, no suffix is needed. TIMESTAMPTZ says “timestamp with timezone” — adding _epoch_ms is redundant. Suffixes are for generic types (BIGINT, INTEGER, TEXT) where the type alone is ambiguous.
CREATE TABLE events (
id TEXT PRIMARY KEY,
created_at TIMESTAMPTZ NOT NULL, -- type says timestamp, no suffix needed
duration_ms INTEGER, -- INTEGER is ambiguous, suffix needed
payload_bytes INTEGER,
api_key_secret TEXT,
retry_count INTEGER, -- no suffix needed, meaning is obvious
domain TEXT NOT NULL
);
| Column | Type | Suffix needed? | Why |
|---|---|---|---|
created_at | TIMESTAMPTZ | no | type encodes semantics |
duration_ms | INTEGER | yes | 142 what? ms vs s vs μs |
payload_bytes | INTEGER | yes | bytes vs KiB vs count |
api_key_secret | TEXT | yes | enables auto-redaction |
retry_count | INTEGER | no | meaning obvious from name |
expires_at | TIMESTAMPTZ | no | type encodes semantics |
cached_epoch_ms | BIGINT | yes | bare integer needs unit |
ORM / struct mapping: Keep the suffix in the struct field name. The suffix is part of the semantic name, not a display concern:
struct Event {
created_at: DateTime<Utc>, // native type — no suffix
duration_ms: i64, // integer — suffix preserves semantics
// duration: i64, // bad — 64-bit what? seconds? ms?
}
Queries: Column aliases in views or query results should also follow AFDATA naming:
SELECT
duration_ms,
payload_bytes,
(cost_input_msats + cost_output_msats) AS total_cost_msats
FROM requests;
Part 2: Output Processing
Transform JSON values for CLI/log output with suffix-driven formatting and automatic secret protection. This applies to any JSON data, regardless of structure.
Two Output Paths
Path 1: Raw JSON Serialization
Return JSON values directly (for example via framework serializer or serde_json::to_string).
No output processing. Values are serialized as-is:
{"user_id": 123, "api_key_secret": "sk-1234567890abcdef", "balance_msats": 50000}
Path 2: CLI / Logs
Format JSON values for terminal/log display.
Automatic processing: secret redaction always; suffix formatting only for Plain.
Input:
{"user_id": 123, "api_key_secret": "sk-1234567890abcdef", "balance_msats": 50000}
JSON: {"api_key_secret":"***","balance_msats":50000,"user_id":123}
YAML:
---
api_key_secret: "***"
balance_msats: 50000
user_id: 123
Plain: api_key=*** balance=50000msats user_id=123
Output Formats
CLI tools should support multiple output formats:
--output json|yaml|plain
--log startup,request,progress,retry,redirect
--verbose
Default is tool-defined. Human-facing interactive use commonly defaults to plain (human-formatted values); scripting/automation contexts use json or yaml (both structure-preserving).
JSON is the canonical format. YAML mirrors it exactly — same keys, same values, different syntax. Plain derives a lossy, human-formatted view from it.
All CLI output formats automatically redact _secret fields. Matching recognizes _secret and _SECRET only. Any _secret value — scalar, object, or array — is replaced with ***. Legacy field names can be protected by passing OutputOptions.redaction.secret_names at serialization time; this opt-in list is exact field-name equality. The Raw output style disables Plain’s formatting suffix stripping while keeping the selected redaction policy; YAML is always structure-preserving regardless of OutputStyle.
One policy, three paths, and a scope that only exists on one of them. The redaction policy governs structured values, argv, and standalone URLs alike. Off disables all three. TraceOnly narrows redaction to the trace subtree of a structured value, and narrows nothing else: a command line and a bare URL have no non-trace half to leave alone, so both are redacted in full, exactly as under the default. TraceOnly says where to redact inside a value, never how much to redact overall — reading it as “off” for argv and URLs would turn a scoping option into a silent opt-out on the two paths that exist to make diagnostics safe.
Stripping the _secret suffix is the companion of having redacted the value, never an unconditional rename. Plain drops the suffix only from a field whose value is the *** marker. Where redaction did not apply — Off, or a field outside trace under TraceOnly — the suffix stays, because removing it would emit the live secret and delete the only mark saying it is one, leaving nothing downstream to protect. So Off renders api_key_secret=sk-live-xxx, not api_key=sk-live-xxx. This makes the registry’s strip_key: true for _secret a statement about the default policy, under which every such value is redacted.
Format characteristics:
- JSON — single-line, original keys, raw values, no sorting (machine-readable), secrets redacted
- YAML — multi-line, original keys, raw values (same semantics as JSON), keys sorted, secrets redacted by default
- Plain — single-line logfmt, human-readable, formatting suffixes stripped, values formatted, secrets redacted by default
yaml
Each JSON line becomes a YAML document, separated by ---. Strings always quoted to avoid YAML pitfalls (no → false, 3.0 → float). YAML is structure-preserving, like JSON: it keeps every key and value exactly as written — no suffix stripping, no value reformatting — regardless of PlainStyle. Secrets automatically redacted.
---
kind: "log"
log:
args:
config_path: "config.yml"
config:
api_key_secret: "***"
dns_ttl_s: 3600
event: "startup"
---
kind: "result"
result:
hash: "abc123"
size_bytes: 456789
trace:
cost_msats: 2056
duration_ms: 1280
plain
Single-line logfmt style. In the default readable style, formatting suffixes are stripped from keys. Secrets automatically redacted.
- Nested keys use dot notation:
trace.duration=1.28s - Values containing ASCII space, tab, newline, carriage return, form feed, vertical tab, NBSP,
=,", or\are quoted;\,", newline, carriage return, tab, form feed, and vertical tab are escaped so each record stays one physical line - Arrays are comma-joined:
fields=email,age - Null values are empty:
RUST_LOG=
kind=log log.args.config_path=config.yml log.config.api_key=*** log.config.dns_ttl=3600s log.event=startup
kind=result result.hash=abc123 result.size=446.1KiB trace.cost=2056msats trace.duration=1.28s
Suffix processing (plain only)
YAML never strips or reformats any key or value — see the yaml section above. Plain applies two transformations:
1. Key stripping — remove the recognized formatting suffix from the key name. The formatted value already encodes the unit, so the suffix is redundant for human readers.
Algorithm: match the longest known suffix from the list below. Each suffix is recognized in two forms: lowercase (_secret) and uppercase (_SECRET). No other casing is matched. Remove the matched suffix from the key. If no suffix matches, keep the key unchanged. Match order (longest first):
_epoch_ms,_epoch_s,_epoch_ns(compound timestamp suffixes)_usd_cents,_eur_cents,_{code}_cents,_{code}_micro(compound currency suffixes;codeis 3-4 ASCII letters)_rfc3339,_minutes,_hours,_days(multi-char suffixes)_msats,_sats,_bytes,_percent,_secret(single-unit suffixes)_jpy,_ns,_us,_ms,_s(short suffixes, matched last to avoid false positives)
Strict string suffixes (_bcp47, _utc_offset, _rfc3339_date, _rfc3339_time) are not key-stripping suffixes. They keep the field’s format contract visible in readable output.
Collision: if two keys in the same object produce the same stripped key (e.g., response_ms and response_bytes both → response), revert both to their original key AND raw value (no formatting). Redaction happens before this step, so collision fallback can never restore a secret value.
| JSON key | Plain key | Why |
|---|---|---|
duration_ms | duration | value shows 1.28s |
size_bytes | size | value shows 446.1KiB |
created_at_epoch_ms | created_at | value shows 2025-02-07T... |
expires_rfc3339 | expires | value passes through |
api_key_secret | api_key | value shows *** |
cpu_percent | cpu | value shows 85% |
balance_msats | balance | value shows 50000msats |
price_usd_cents | price | value shows $9.99 |
DATABASE_URL_SECRET | DATABASE_URL | uppercase _SECRET matched |
CACHE_TTL_S | CACHE_TTL | uppercase _S matched |
language_bcp47 | language_bcp47 | strict string format, key unchanged |
timezone_utc_offset | timezone_utc_offset | fixed-offset string, key unchanged |
invoice_due_rfc3339_date | invoice_due_rfc3339_date | RFC 3339 full-date string, key unchanged |
market_open_rfc3339_time | market_open_rfc3339_time | RFC 3339 partial-time string, key unchanged |
config_path | config_path | no suffix, unchanged |
user_id | user_id | no suffix, unchanged |
2. Value formatting — transform the value for human readability. Same suffix matching as key stripping (lowercase or uppercase only):
_ns,_us,_ms,_s→ append unit (450000ns,830μs,42ms,3600s)_mswith absolute value ≥ 1000 → convert to seconds (1280→1.28s,-1500→-1.5s)_minutes,_hours,_days→ append unit (30 minutes,24 hours)_epoch_ms/_epoch_s/ decimal-string_epoch_ns→ RFC 3339 (2024-02-14T00:00:00.000Z), negative values produce pre-1970 dates_rfc3339→ pass through_bytes→ human-readable (456789→446.1KiB); negative and fractional byte values fall through as raw values_percent→ append%(85→85%,99.9→99.9%)_msats→ append unit (2056msats)_sats→ append unit (1234sats)_usd_cents→ dollars (999→$9.99,-499→-$4.99)_eur_cents→ euros (850→€8.50,-850→-€8.50)- other
_{code}_cents→ major unit with code (15050→150.50 THB,-15050→-150.50 THB), wherecodeis 3-4 ASCII letters _{code}_micro→ major unit with six decimals and code (170000→0.170000 USD,-170000→-0.170000 USD), wherecodeis 3-4 ASCII letters_jpy→ yen (1500→¥1,500,-1500→-¥1,500)_secret→***(already applied by the redaction phase; the formatter does not perform a second, divergent redaction pass)
Strict string fields such as _bcp47, _utc_offset, _rfc3339_date, and _rfc3339_time are not value-formatting suffixes; their string values pass through unchanged.
A _url field value is preserved byte-for-byte in Plain output except for the redacted secret spans (userinfo password, _secret-suffixed/secret_names query parameters): the _url key is not stripped, and formatting suffixes that appear inside the URL — ?timeout_ms=5000, ?size_bytes=1048576 — are not reformatted (5s, 1.0MiB) or stripped, because the URL must round-trip to its server exactly. URL key-stripping/value-formatting applies to JSON object keys, never to query parameters inside a string value. YAML never strips or reformats any key/value in the first place, so a _url field renders there exactly like any other field. This is pinned by the url_params_redacted_not_reformatted case in spec/fixtures/output_formats.json.
Type constraints: _bytes requires a non-negative integer; _epoch_* and fiat suffixes (_usd_cents, _eur_cents, _jpy, _{code}_cents, _{code}_micro) require integers, with fiat explicitly allowing either sign. Duration, Bitcoin, and _percent suffixes accept any number. When the value type doesn’t match, formatting falls through to the raw value with the original key preserved. An integral-valued float counts as an integer for the integer-required suffixes (3.0 is treated as 3): a JSON number’s value, not its lexical form, decides, because JavaScript cannot distinguish 3 from 3.0 after parsing.
Number rendering: a number is rendered for YAML/plain by the shared fixture-defined decimal form: integral-valued floats drop their trailing .0 (3.0 → 3), exponent markers use lowercase e, and exponent signs/leading zeroes are normalized (1e-07 → 1e-7). Integers beyond 2⁵³ are preserved exactly by Rust, Go, and Python; JavaScript loses precision on them (see the _epoch_ns precision note above).
Key ordering
YAML sorts keys by UTF-16 code unit order (JCS, RFC 8785 §3.2.3) without stripping suffixes. Plain sorts keys the same way, but after stripping. For ASCII keys — the common case — this equals simple byte-order sorting.
In plain logfmt, nested keys are flattened to dot notation before sorting. Sort by the full dot path: args.input_path < code < config.api_key < trace.duration.
JSON output is unordered per the JSON specification. YAML and plain sort for deterministic, cross-language-consistent output.
Using AFDATA Without Part 3
Parts 1 and 2 (naming + output processing) work with any JSON structure — no protocol template needed:
{"user_id": 123, "created_at_epoch_ms": 1738886400000, "balance_msats": 50000000, "api_key_secret": "sk-..."}
Plain: api_key=*** balance=50000000msats created_at=2025-02-07T00:00:00.000Z user_id=123
This works with REST APIs, GraphQL, database results, config files — anywhere you have structured data. Just use AFDATA naming and let output processing handle the rest.
Part 3: Protocol Template (Recommended, Optional)
A recommended structure for program output. This part is optional — adopt it when you want consistent structure across CLI tools, streaming output, or internal protocols.
Core Fields
Required:
kind— protocol discriminator:"result","error","progress", or"log"- a payload field whose name matches
kind
Recommended:
trace— execution context (duration, source, resource usage)
trace, when present, is a JSON object. result, progress, and log
payloads are tool-defined valid JSON values. error is a JSON object with
required non-empty code and message, optional hint, and tool-defined
extension fields.
CLI Event Framing
Structured CLI programs emit complete AFDATA protocol v1 events. Where each
event goes is set by the program’s consumption mode (see Channel policy below);
how each event is framed depends on --output:
- JSON multi-event output is JSONL/NDJSON: one complete event per line.
- Plain multi-event output is one display event per line.
- YAML multi-event output uses an explicit
---document boundary for every event.
Agent-facing machine input remains JSON. YAML and plain are display formats.
Channel policy — the stream an event lands on follows the consumption mode, not the event’s shape. There are two modes:
Finite one-shot commands (the default). A command that runs, emits at most
one terminal event, and exits. Its events split by kind, following Unix
stream conventions:
kind:"result"(and any data/findings payload the caller captures) →stdoutkind:"error"→stderrkind:"progress"andkind:"log"→stderr(they are diagnostics)- routing follows
kind, not the exit code: akind:"result"payload always goes tostdouteven when the command exits non-zero (e.g. a search that completed but matched nothing), andkind:"error"always goes tostderr stdouttherefore carries only successful payloads, sox=$(tool …)never captures an error as data,tool … >/dev/nullnever swallows a diagnostic, andtool … | nextnever pipes an error envelope in as input
Event streams. A process whose consumer reads interleaved
log/progress/result/error events in order (a long-running or agentic
producer). Every event stays on ONE stream so ordering is preserved; splitting a
stream across stdout and stderr would lose ordering and is prohibited. That
stream’s destination is chosen by the emitter (stdout by default).
A multiplex transport owns lifecycle state per request or logical stream.
CliEmitter represents one such logical stream; AFDATA does not provide a
global keyed-emitter registry or treat the whole multiplex connection as one
finite CLI invocation.
Choosing the mode. A command is an event stream when it produces more than one caller-needed output over time; otherwise it is finite. Two shapes qualify even though each looks like a single operation:
- a payload delivered in chunks — an opening event, N batches carrying the rows, a terminator. The batches carry the data, which therefore does not fit in the single terminal event finite mode allows.
- an operation that reports something the caller must act on before it can finish — a served address to connect to, a URL a human must visit — and then reports its outcome. Both are caller-needed, and an unbounded wait separates them.
Neither fits finite mode, whose defining limit is at most one terminal event.
Splitting either across stdout and stderr strands the caller’s data on the
diagnostic stream: the row batches or the address land on stderr while only a
terminator or a shutdown acknowledgement reaches stdout.
An event-stream command therefore defaults to --output-to stdout rather than
split, and rejects an explicit --output-to split as a usage error — the same
way a raw-scalar reader rejects a non-default destination. Its intermediate
data-carrying events stay kind:"progress": once the whole stream shares one
destination, kind no longer decides routing, and the command still emits
exactly one terminal event.
The invariant this preserves: data the caller captures is never routed to the
diagnostic stream. A kind:"progress" event carrying a payload the caller
must read is not a finite command with an unusual payload; it is an event stream
that has not declared itself.
Destination selection:
- CLI tools SHOULD expose
--output-to <split|stdout|stderr>, defaultsplitfor a finite command andstdoutfor an event-stream command.splitis finite mode above.stdout/stderrselect event-stream mode and collapse the whole envelope stream — everykind, includingerror— onto that one stream.--outputselects an event’s format;--output-toselects its destination; the two are orthogonal. --output-to stdoutrestores the “read one stream, branch onkind” contract for a consumer that wants it;--output-to split(the default) is safe for shell capture and pipelines.- Raw-scalar reader commands (a command whose success output is an unwrapped
value, not an envelope) are intrinsically
splitand reject a non-default--output-toas a usage error.
In finite mode stderr carries both the formatted AFDATA error/progress/
log envelopes and native diagnostics (Rust panics, Python tracebacks). A
consumer reading the error stream parses AFDATA JSON lines and tolerates
interleaved native, non-JSON bytes.
Optional stream redirection:
- CLI tools and services MAY expose
--stdout-file <PATH>and--stderr-file <PATH> - unset file flags leave the corresponding stream unchanged
- when enabled, stdout bytes are appended to the
--stdout-filepath instead of the original stdout destination - when enabled, stderr bytes are appended to the
--stderr-filepath instead of the original stderr destination - an event stream can be sent to a file by collapsing it (
--output-to stdout) and redirecting that stream (--stdout-file <PATH>); there is no separate events-file flag --outputcontinues to select stdout format (json,yaml,plain); it does not select stream destinations- implementations SHOULD install stream redirection before version/help handling, logging/tracing initialization, and other early output
- startup failures to create/open the files SHOULD fail startup with a structured error when a stream is still available
- redirection is a process-level file-descriptor concern applied beneath the
logical channel routing above; native panics/tracebacks stay raw
stderrbytes - no application-level rotation is implied; rotate with external tooling
- this is stream redirection, not a second AFDATA protocol channel and not stream copying
Recommended enforcement:
- every
stderrwrite goes through the emitter’s formatted diagnostic/error path; no ad-hocstderrwriting that bypasses the formatter in runtime code - Rust: route through the emitter; keep native panic output as the only
unstructured
stderr(no strayeprintln!/std::io::stderrin runtime code) - Go/Python/TypeScript: source-policy tests or lint rules that fail on ad-hoc stderr APIs in runtime code, with the emitter’s own sink as the sanctioned exception
Finite structured CLI event streams follow:
(log | progress)* -> exactly one (result | error) -> end
Log and progress payloads are tool-defined JSON values with no required or reserved payload fields. Traditional logging adapters commonly add message and level; progress producers may add a human-readable message. These are conventions, not protocol requirements. Projects that need timestamps add timestamp_epoch_ms explicitly.
Log fields are redacted by field name at emit time — the same _secret/_url rule as all other output, applied by the formatter, not by scanning rendered values. Emit secrets as named fields (api_key_secret) so the rule can see them. Logging a whole object pre-rendered to a single string (e.g. a language’s debug/inspect form) defeats redaction, because the inner field names are no longer visible: build a structured value and redact it before logging instead.
The top-level kind values are reserved: log, progress, result, and error. Tool-defined codes belong inside the corresponding payload.
Error payload codes: Use specific codes instead of generic "error":
"not_found","unauthorized","validation_error","rate_limit","internal_error", etc.- Generic
"error"is supported but specific codes are preferred
Progress and log payloads may add tool-defined fields such as event: "request" or phase: "sync".
Not all phases are required. A simple CLI tool may emit only a result line. A long-running service may never emit a result.
Startup Diagnostic Event
kind: "log" with log.event: "startup". Optional. Emitted once at the beginning if diagnostic logging is enabled.
{"kind":"log","log":{"message":"startup","level":"info","event":"startup","version":"0.1.0","argv":["tool","--log","startup"],"config":{"api_key_secret":"***","dns_ttl_s":3600},"args":{"config_path":"config.yml"},"env":{"RUST_LOG":null,"DATABASE_URL_SECRET":"***"}},"trace":{}}
Startup payload fields are tool-defined. Common fields:
version— tool version stringargv— raw CLI argv arrayconfig— resolved configuration (recommended)args— parsed CLI arguments (optional)env— environment variables the program reads (nullif unset, optional)
Status
kind is the protocol discriminator. progress and log payload content is tool-defined. Include trace for execution context when it helps debugging.
{"kind": "progress", "progress": {"current": 3, "total": 10, "message": "indexing spores"}, "trace": {"duration_ms": 500}}
{"kind":"log","log":{"message":"POST /v1/chat completed","level":"info","event":"request","method":"POST","path":"/v1/chat","http_status":200},"trace":{"latency_ms":42}}
Result
kind:"result" MUST be emitted only when the command intent was completely
fulfilled. Any incomplete fulfillment, including partial completion, MUST emit
kind:"error". An agent watching a finite stream can treat either as the
unique terminal event.
Always include trace for execution context — duration, data sources, resource usage, query details.
Success:
{"kind": "result", "result": {"hash": "abc123", "size_bytes": 456789}, "trace": {"duration_ms": 1280, "tokens_input": 512}}
Error:
Simple message:
{"kind": "error", "error": {"code": "config_not_found", "message": "config file not found", "retryable": false}, "trace": {"duration_ms": 3}}
With actionable hint:
{"kind": "error", "error": {"code": "connection_refused", "message": "connection refused", "retryable": false, "hint": "check --host/--port or PGHOST/PGPORT environment variables"}, "trace": {"duration_ms": 3}}
The hint field is optional. When present, it provides an actionable suggestion for the user or agent to resolve the error. Omit hint when no specific remediation is available.
Error details are direct extension fields inside the error payload:
{"kind": "error", "error": {"code": "not_found", "message": "user not found", "retryable": false, "resource": "user", "id": 123}, "trace": {"duration_ms": 8}}
More examples:
{"kind": "error", "error": {"code": "validation_error", "message": "invalid fields", "retryable": false, "fields": ["email", "age"]}, "trace": {"duration_ms": 2}}
{"kind": "error", "error": {"code": "unauthorized", "message": "invalid token", "retryable": false}, "trace": {"duration_ms": 5}}
{"kind": "error", "error": {"code": "rate_limit", "message": "rate limited", "retryable": false, "retry_after_s": 60, "quota_remaining": 0}, "trace": {"duration_ms": 1}}
Best Practices
Always include trace field. Even simple operations should report execution context:
duration_ms— operation durationsource— data source (db, cache, api, file)- Resource usage —
tokens_input,tokens_output,cost_msats,memory_bytes - Metadata —
query,method,path,model
Good (with trace):
{"kind": "result", "result": {"count": 42}, "trace": {"duration_ms": 150, "source": "db"}}
{"kind": "error", "error": {"code": "not_found", "message": "not found", "retryable": false}, "trace": {"duration_ms": 5}}
Also good:
{"kind": "result", "result": {"count": 42}, "trace": {"duration_ms": 150, "source": "db"}}
{"kind": "error", "error": {"code": "validation_error", "message": "invalid input", "retryable": false, "fields": [...]}, "trace": {"duration_ms": 2}}
Avoid (missing trace):
{"kind": "result", "result": {"count": 42}}
{"kind": "error", "error": {"code": "not_found", "message": "not found", "retryable": false}}
Missing trace makes debugging harder. Agents can’t analyze performance, cost, or data flow without execution context.
Validation profiles
The validate_protocol_event(event, strict) / validate_protocol_stream(events, strict) APIs enforce protocol compliance. With strict=false, only mandatory MUST rules are enforced: envelope shape, error payload requirements, and finite-stream lifecycle. With strict=true (the default in Python/TS), additional recommendations are required:
- every event includes an object-valued
trace - every error payload includes
retryableas a boolean
Passing the base validator proves mandatory conformance only. Use the strict validator when claiming conformance with these recommendations. Structured version metadata may intentionally use the base profile when no execution trace exists.
Agent consumption
- Read
kindon every line and read the same-named payload (event[kind]). kind:"log"withlog.event:"startup"describes resolved startup configuration.kind:"result"orkind:"error"completes a finite operation.kind:"log"andkind:"progress"are non-terminal events before that terminal event.
Usage in HTTP Services
The protocol structure can be used in REST APIs. Choose output path explicitly:
- raw JSON serialization for untouched payloads
- formatter output (
json|yaml|plain) when redaction/formatting is required
REST API Examples
Response body follows the protocol structure:
HTTP 200:
{"kind": "result", "result": {"balance_msats": 97900}, "trace": {"source": "redb", "duration_ms": 3}}
HTTP 404:
{"kind": "error", "error": {"code": "not_found", "message": "user not found", "retryable": false, "resource": "user", "id": 123}, "trace": {"duration_ms": 5}}
HTTP 402:
{"kind": "error", "error": {"code": "insufficient_balance", "message": "insufficient balance", "retryable": false, "balance_msats": 0, "required_msats": 2056}, "trace": {"source": "redb", "duration_ms": 2}}
MCP Tool Response
Same structure, raw JSON:
{"kind": "result", "result": {"files": ["src/main.rs"]}, "trace": {"source": "glob", "matched": 1, "duration_ms": 12}}
Streaming (SSE)
JSONL stream, raw JSON per line:
{"kind":"log","log":{"message":"startup","level":"info","event":"startup","config":{"model":"gpt-4","max_tokens":1024},"args":{},"env":{}},"trace":{}}
{"kind": "progress", "progress": {"current": 1, "total": 5, "message": "processing"}, "trace": {"duration_ms": 500}}
{"kind": "result", "result": {"answer": "..."}, "trace": {"tokens_input": 512, "duration_ms": 1280}}
One Protocol, Multiple Contexts
| Context | Output | Secret Protection |
|---|---|---|
| CLI / Logs | JSONL, one-line plain events, or YAML documents | ✅ Automatic |
| HTTP body (raw path) | JSON body (raw Value) | Use redacted_value before framework serialization |
| MCP tool (raw path) | JSON (raw Value) | Use redacted_value before SDK serialization |
| SSE stream (raw path) | JSONL (raw JSON) | Use redacted_value before emitting events |
All contexts can use the protocol structure from Part 3. kind, its matching
payload field, and optional object-valued trace are standardized. CLI/logs
apply output formatting and secret protection from Part 2. Raw-path serializers
return JSON values unchanged unless the program explicitly calls
redacted_value. For an event stream (interleaved events consumed in
order), keep every event on one stream and do not split across stdout and
stderr — splitting loses ordering. A finite one-shot command instead
splits by kind (result → stdout, error → stderr); see CLI Event
Framing for both modes and the --output-to selector.
Complete Example: CLI Tool
A complete example showing all three parts working together. A backup tool that uploads files to cloud storage.
CLI Invocation
cloudback --api-key-secret sk-1234567890abcdef --timeout-s 30 --max-file-size-bytes 10737418240 /data/backup.tar.gz
Flag names use AFDATA suffixes in kebab-case. An agent reading --help knows --timeout-s is seconds and --api-key-secret should be redacted — no documentation needed.
Raw JSON (before output processing)
The tool converts CLI flags from kebab-case to snake_case and emits a startup diagnostic event when enabled:
{
"kind": "log",
"log": {
"timestamp_epoch_ms": 1710000000000,
"message": "startup",
"level": "info",
"event": "startup",
"config": {
"api_key_secret": "sk-1234567890abcdef",
"endpoint": "https://storage.example.com",
"timeout_s": 30,
"max_file_size_bytes": 10737418240
},
"args": {
"input_path": "/data/backup.tar.gz",
"compression_level": 9
}
},
"trace": {}
}
Field names encode semantics:
api_key_secret→ agent knows to redacttimeout_s→ 30 secondsmax_file_size_bytes→ 10GiB in bytes
Output Formats (Part 2: Output Processing)
JSON (raw, for machines):
{"kind":"log","log":{"message":"startup","level":"info","event":"startup","config":{"api_key_secret":"***","endpoint":"https://storage.example.com","timeout_s":30,"max_file_size_bytes":10737418240},"args":{"input_path":"/data/backup.tar.gz","compression_level":9}},"trace":{}}
YAML (structured, original keys/values — same semantics as JSON, machine-readable):
---
kind: "log"
log:
args:
compression_level: 9
input_path: "/data/backup.tar.gz"
config:
api_key_secret: "***"
endpoint: "https://storage.example.com"
max_file_size_bytes: 10737418240
timeout_s: 30
event: "startup"
level: "info"
message: "startup"
trace: {}
Plain (single-line logfmt, formatting suffixes stripped, for compact scanning):
kind=log log.args.compression_level=9 log.args.input_path=/data/backup.tar.gz log.config.api_key=*** log.config.endpoint=https://storage.example.com log.config.max_file_size=10.0GiB log.config.timeout=30s log.event=startup log.level=info log.message=startup
Note:
- Key stripping (Plain only): formatting suffixes such as
api_key_secret→api_key,timeout_s→timeout,max_file_size_bytes→max_file_size(_secretonly once its value has been redacted — see above) - Secret protection:
api_key_secretredacted in all three formats - Suffix formatting (Plain only):
_bytes→10.0GiB,_s→30s; JSON and YAML keep the raw integer
Progress Update (Part 3: Protocol Template)
{"kind":"progress","progress":{"current":3,"total":10,"message":"uploading chunks"},"trace":{"duration_ms":5420,"uploaded_bytes":3221225472}}
YAML:
---
kind: "progress"
progress:
current: 3
message: "uploading chunks"
total: 10
trace:
duration_ms: 5420
uploaded_bytes: 3221225472
Plain:
kind=progress progress.current=3 progress.message="uploading chunks" progress.total=10 trace.duration=5.42s trace.uploaded=3.0GiB
Final Result
{"kind": "result", "result": {"backup_url": "https://storage.example.com/backup.tar.gz", "size_bytes": 10485760, "checksum": "sha256:abc123...", "uploaded_at_epoch_ms": 1738886400000}, "trace": {"duration_ms": 15300, "chunks": 10, "retries": 2}}
YAML:
---
kind: "result"
result:
backup_url: "https://storage.example.com/backup.tar.gz"
checksum: "sha256:abc123..."
size_bytes: 10485760
uploaded_at_epoch_ms: 1738886400000
trace:
chunks: 10
duration_ms: 15300
retries: 2
Plain:
kind=result result.backup_url=https://storage.example.com/backup.tar.gz result.checksum=sha256:abc123... result.size=10.0MiB result.uploaded_at=2025-02-07T00:00:00.000Z trace.chunks=10 trace.duration=15.3s trace.retries=2
What This Demonstrates
-
Part 1 (Naming): Every field is self-describing — from CLI flags (
--timeout-s,--api-key-secret) to JSON fields (timeout_s,uploaded_at_epoch_ms). Same suffixes, same semantics, kebab↔snake mapping -
Part 2 (Output Processing): Three formats for different needs
- JSON: single-line, original keys, raw values, for programs and logs
- YAML: multi-line, original keys, raw values (same semantics as JSON), for machine-readable multi-line output
- Plain: single-line logfmt, formatting suffixes stripped, values formatted, for compact scanning
- All formats protect secrets automatically
-
Part 3 (Protocol): Consistent structure across all output —
kindidentifies the event type and its same-named payload,traceprovides execution context, and payload fields remain tool-defined
Key insight: The same naming convention flows from CLI flag (--timeout-s 30) to JSON/YAML field (timeout_s: 30) to Plain’s human-formatted output (timeout: 30s). An agent reading --help, JSON, or YAML sees the same self-describing field name; Plain trades the field name for a human-formatted value. No documentation needed at any layer.