Commit Graph

3 Commits

Author SHA1 Message Date
Gaurav Bajpai
67d919947f
fix: address PRizm round-2 review findings from internal port @W-22196528@
Ports the 14 applicable round-2 fixes from the internal PRizm review to the
external afv-library source. (plugin.json finding #13 is internal-only and
does not apply here.)

Accepted fixes:
  1  bdt_analyze.py select_definition: isinstance(index, int) -> type(index) is int
     (reject bool subclass of int).
  2  bdt_analyze.py cmd_formula: precompute upstream fields_produced map to
     eliminate O(consumed x upstream) recomputation.
  3  bdt_analyze.py topo_order: replace sort-on-every-iteration with deque-based
     Kahn (preserves deterministic order).
  4  test_bdt_analyze.py test_all_subcommands_work_on_api_input: wrap loop in
     self.subTest for per-iteration failure reporting.
  5  test_bdt_analyze.py test_invalid_json_raises: register addCleanup BEFORE
     write_text so cleanup runs even if write fails.
  6  SKILL.md: remove misleading "strip __c suffix" field-trace troubleshooting
     advice; replace with source-DMO passthrough guidance.
  7  SKILL.md: rewrite cycle error message to direct user to Data Cloud viewer
     (re-export will not fix a cycle in the BDT definition).
  8  SKILL.md: add explicit --definition N flag documentation and list every
     subcommand that accepts it.
  9  SKILL.md: refine I2 routing guidance — pick the earliest formula/
     computeRelative node, not the output mapping.
  11 bdt-function-catalog.md: clean floor() wording ("always rounds down on the
     number line").
  12 bdt-node-catalog.md: businessType enum — assert list is canonical/complete;
     instruct skill to surface+flag any unknown value.
  14 append_and_split.json: change split subject from fragile
     CustomerFullName__c split-on-space to ChannelOrderKey__c split-on-hyphen;
     avoids name-parsing pitfall (middle names / multi-space).
  15 window_and_aggregate.json: use the computed OrderRank__c via a new
     FIRST_ORDER_AMOUNT formula node (case when rank=1 then amount else 0)
     feeding AGG_BY_ACCOUNT; illustrates the rank+formula+aggregate idiom.
  16 joins_and_filters.json: add explicit IS_NOT_NULL filter expression on
     GrandTotalAmount__c (defense-in-depth beyond GREATER_THAN 0).

Rejected finding (rationale posted as PR comment):
  10 bdt-node-catalog.md split section — reviewer claimed split is a
     pipeline-branching node. Canonical sources (core-262
     SplitParametersInputRepresentation.java + Salesforce help DITA
     c360_a_batch_transform_split.xml) both describe split as a
     string-splitting operation. Current docs are correct; not changing.

Tests: 92/92 passing.
2026-04-24 10:55:18 +05:30
Gaurav Bajpai
db1aca7221
fix: address PRizm review findings from internal port PR @W-22196528@
Addresses 4 critical findings surfaced during PRizm code review of the
internal plugin port of this skill (internal PR #19). Applying the same
fixes here keeps the external canonical source and the internal port in
sync.

1. Test cleanup discipline: switch `test_invalid_json_raises` from
   `try/finally` to `self.addCleanup(p.unlink, missing_ok=True)` — the
   unittest-idiomatic way to guarantee temp-file cleanup regardless of
   how the test exits.

2. `floor()` description in bdt-function-catalog.md: the old row was
   self-contradictory ("toward zero" AND "toward next integer up" in
   the same cell). Replace with a single coherent definition: rounds
   toward negative infinity; for negatives rounds away from zero
   (e.g., `floor(-2.3) = -3`).

3. Split-node documentation in bdt-node-catalog.md: the old doc claimed
   `split` routes rows into downstream branches via `branches[]` with
   per-branch predicates. That is not the canonical schema. Per
   `SplitParametersInputRepresentation` in core-262-public, `split` is
   a string-splitting operation: one `sourceField` + `delimiter` →
   N `targetFields` (one row in, one row out; columns added). Rewrote
   the section with the correct parameters, lineage effect, gotchas,
   and a canonical example. Row-routing belongs in `filter` nodes.

4. Sample `assets/sample_bdts/append_and_split.json`: the old sample
   used the invented `branches[]` shape AND routed the same split into
   two downstream outputs that each expected different rows — which is
   not how `split` works. Rewrote the sample so:
     - `appendV2` unions two order sources (unchanged intent).
     - `split` uses canonical `{sourceField, delimiter, targetFields}`
       splitting `CustomerFullName__c` into first + last name columns.
     - One downstream output consumes the new columns (removes the
       fake two-branch fan-out).

Tests: 92/92 passing. Sample parses and runs through `bdt_analyze.py
summary` cleanly (5 nodes: 2 load + 1 appendV2 + 1 split + 1 outputD360).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-24 10:19:36 +05:30
Gaurav Bajpai
fc4d14bce0
test(bdt): unit tests + adversarial fixtures @W-22196528@
Adds the full test suite for bdt_analyze.py: 92 unit tests backed by
9 JSON fixtures (4 normative + 5 adversarial). All tests stdlib-only,
runnable via `python3 -m unittest tests.test_bdt_analyze`.

tests/test_bdt_analyze.py coverage:

DAG primitives:
- roots() / sinks() / topo_order() on minimal + branching graphs.
- upstream() / downstream() reachability correctness.
- Cycle tolerance: topo_order + traversals do not infinite-loop or
  raise when the graph contains a cycle; cycle is reported via
  broken_references-adjacent signals.
- Broken-reference tolerance: nodes referencing non-existent upstreams
  are handled gracefully; broken_references() enumerates them.

CLI subcommands (happy-path + error-path for each):
- summary, sources, outputs, stages, nodes, node, lineage,
  field-trace, formula, definitions — each exercised in both
  human-readable and --json modes.
- Exit-code discipline: unknown node -> exit 2, malformed input ->
  exit 3, success -> exit 0.
- Unknown-action graceful degradation: nodes whose action isn't in
  the catalog still appear in summary/nodes output with a generic
  label rather than crashing.
- Output size budgets: default-mode outputs are asserted to stay
  within configured character caps.

Field discovery:
- fields_produced() / fields_consumed() / _scrape_field_refs()
  heuristics tested against both clean and noisy formula bodies.
- TestFieldTraceFocus — narrowing tests preventing false-positive
  substring hits in formula text.

Dual-shape input:
- TestInputShapeDetection — editor export vs. Connect API
  single-definition auto-detection.
- TestMultiDefinition — Connect API multi-definition payload:
  definitions subcommand, --definition N routing, out-of-range
  index error handling.

Fixtures (tests/fixtures/):
- minimal.json               — smallest valid editor-export BDT.
- window_and_aggregate.json  — window + aggregate composition.
- api_input_single.json      — Connect API single-definition payload.
- api_input_multi.json       — Connect API multi-definition payload.

Adversarial fixtures (tests/fixtures/adversarial/):
- cycle.json           — graph with a cycle; parser must not hang.
- broken_ref.json      — node referencing a non-existent upstream.
- unknown_action.json  — node with an action not in the catalog.
- no_ui.json           — BDT missing the UI layer entirely.
- empty_nodes.json     — valid envelope but zero nodes.

Suite status: Ran 92 tests, OK.

@W-22196528@
2026-04-23 23:29:05 +05:30