Addresses 4 critical findings surfaced during PRizm code review of the
internal plugin port of this skill (internal PR #19). Applying the same
fixes here keeps the external canonical source and the internal port in
sync.
1. Test cleanup discipline: switch `test_invalid_json_raises` from
`try/finally` to `self.addCleanup(p.unlink, missing_ok=True)` — the
unittest-idiomatic way to guarantee temp-file cleanup regardless of
how the test exits.
2. `floor()` description in bdt-function-catalog.md: the old row was
self-contradictory ("toward zero" AND "toward next integer up" in
the same cell). Replace with a single coherent definition: rounds
toward negative infinity; for negatives rounds away from zero
(e.g., `floor(-2.3) = -3`).
3. Split-node documentation in bdt-node-catalog.md: the old doc claimed
`split` routes rows into downstream branches via `branches[]` with
per-branch predicates. That is not the canonical schema. Per
`SplitParametersInputRepresentation` in core-262-public, `split` is
a string-splitting operation: one `sourceField` + `delimiter` →
N `targetFields` (one row in, one row out; columns added). Rewrote
the section with the correct parameters, lineage effect, gotchas,
and a canonical example. Row-routing belongs in `filter` nodes.
4. Sample `assets/sample_bdts/append_and_split.json`: the old sample
used the invented `branches[]` shape AND routed the same split into
two downstream outputs that each expected different rows — which is
not how `split` works. Rewrote the sample so:
- `appendV2` unions two order sources (unchanged intent).
- `split` uses canonical `{sourceField, delimiter, targetFields}`
splitting `CustomerFullName__c` into first + last name columns.
- One downstream output consumes the new columns (removes the
fake two-branch fan-out).
Tests: 92/92 passing. Sample parses and runs through `bdt_analyze.py
summary` cleanly (5 nodes: 2 load + 1 appendV2 + 1 split + 1 outputD360).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds the curated reference library the skill loads on demand, plus
four synthetic sample BDTs used by docs, tests, and LLM-mode demos.
references/ (4 curated Markdown files):
- bdt-reference.md — top-level BDT JSON anatomy: envelope,
nodes, edges, UI layer, definitions,
businessType semantics. Cites the core-262
upstream JSON schema and Connect API spec.
- bdt-node-catalog.md — every node type (DMO Source, DMO Sink,
Filter, Join, Union, Aggregate, Window,
Formula, Split, Append, etc.) with its
required/optional fields and typical
usage. Audited against core-262 enums.
- bdt-function-catalog.md — the expression-language function surface
(string, numeric, date, conditional,
aggregate). Grouped by category with
signature + one-line semantics.
- bdt-window-functions.md — windowing operators (ROW_NUMBER, RANK,
LEAD/LAG, running aggregates) with PARTITION
BY / ORDER BY grammar and gotchas.
assets/sample_bdts/ (4 synthetic, dependency-free BDTs):
- minimal_dmo_to_dmo.json — smallest valid BDT (1 source, 1 sink).
- joins_and_filters.json — join + filter composition.
- window_and_aggregate.json — window function + aggregate in one graph.
- append_and_split.json — append-then-split branching topology.
Grounding rules enforced in this commit:
- Every claim in references/ cites an upstream source (core-262 JSON
schema, Connect API reference, or the Data Cloud BDT editor spec).
No speculative content.
- No raw DITA or internal-only documentation is shipped; references
are synthesized from public-facing material.
- BusinessTypeEnum values use the canonical camelCase casing from
core-262 (case-cleanup fix included here).
- Sample BDTs are original synthetic fixtures, not redacted customer
data. Each is small enough to read end-to-end.
@W-22196528@