Commit Graph

2 Commits

Author SHA1 Message Date
Gaurav Bajpai
67d919947f
fix: address PRizm round-2 review findings from internal port @W-22196528@
Ports the 14 applicable round-2 fixes from the internal PRizm review to the
external afv-library source. (plugin.json finding #13 is internal-only and
does not apply here.)

Accepted fixes:
  1  bdt_analyze.py select_definition: isinstance(index, int) -> type(index) is int
     (reject bool subclass of int).
  2  bdt_analyze.py cmd_formula: precompute upstream fields_produced map to
     eliminate O(consumed x upstream) recomputation.
  3  bdt_analyze.py topo_order: replace sort-on-every-iteration with deque-based
     Kahn (preserves deterministic order).
  4  test_bdt_analyze.py test_all_subcommands_work_on_api_input: wrap loop in
     self.subTest for per-iteration failure reporting.
  5  test_bdt_analyze.py test_invalid_json_raises: register addCleanup BEFORE
     write_text so cleanup runs even if write fails.
  6  SKILL.md: remove misleading "strip __c suffix" field-trace troubleshooting
     advice; replace with source-DMO passthrough guidance.
  7  SKILL.md: rewrite cycle error message to direct user to Data Cloud viewer
     (re-export will not fix a cycle in the BDT definition).
  8  SKILL.md: add explicit --definition N flag documentation and list every
     subcommand that accepts it.
  9  SKILL.md: refine I2 routing guidance — pick the earliest formula/
     computeRelative node, not the output mapping.
  11 bdt-function-catalog.md: clean floor() wording ("always rounds down on the
     number line").
  12 bdt-node-catalog.md: businessType enum — assert list is canonical/complete;
     instruct skill to surface+flag any unknown value.
  14 append_and_split.json: change split subject from fragile
     CustomerFullName__c split-on-space to ChannelOrderKey__c split-on-hyphen;
     avoids name-parsing pitfall (middle names / multi-space).
  15 window_and_aggregate.json: use the computed OrderRank__c via a new
     FIRST_ORDER_AMOUNT formula node (case when rank=1 then amount else 0)
     feeding AGG_BY_ACCOUNT; illustrates the rank+formula+aggregate idiom.
  16 joins_and_filters.json: add explicit IS_NOT_NULL filter expression on
     GrandTotalAmount__c (defense-in-depth beyond GREATER_THAN 0).

Rejected finding (rationale posted as PR comment):
  10 bdt-node-catalog.md split section — reviewer claimed split is a
     pipeline-branching node. Canonical sources (core-262
     SplitParametersInputRepresentation.java + Salesforce help DITA
     c360_a_batch_transform_split.xml) both describe split as a
     string-splitting operation. Current docs are correct; not changing.

Tests: 92/92 passing.
2026-04-24 10:55:18 +05:30
Gaurav Bajpai
9f1603ed2f
feat(bdt): Python parser core + CLI subcommands @W-22196528@
Adds scripts/bdt_analyze.py — a generic, stdlib-only DAG parser and
query CLI for Salesforce Data Cloud BDT JSON. 1,379 lines.

Parser primitives:
- DataTransform.from_path / from_dict — accepts three input shapes:
  editor export ({version, nodes, ui, ...}), Connect API
  single-definition ({name, label, type, definition: {...}}), and
  Connect API multi-definition ({name, definitions: [{name, label,
  type, definition}, ...]}).
- Node dataclass with ui_label / ui_description fallback resolution.
- roots() / sinks() / topo_order() — Kahn's algorithm with
  deterministic tie-break for reproducible output.
- upstream() / downstream() traversal resilient to broken refs and
  cycles (does not infinite-loop on self-edges or cycles).
- broken_references(), fields_produced(), fields_consumed(),
  _scrape_field_refs() heuristics for field-trace discovery.

CLI (argparse, 10 subcommands, each supports --json for machine
output):
- summary      — node counts, source/sink counts, stage totals.
- sources      — list source nodes (no upstream).
- outputs      — list sink nodes (no downstream).
- stages       — topologically ordered stages.
- nodes        — flat node listing with labels.
- node <name>  — per-node detail (action, inputs, outputs, fields,
                 UI label/description).
- lineage <node>       — upstream + downstream chain from a node.
- field-trace <field>  — which nodes produce/consume a given field.
- formula <node>       — extract formulas/expressions from a node.
- definitions          — lists definitions in multi-definition
                         payloads; every other subcommand accepts
                         --definition N (default 0) to route into a
                         specific definition within the payload.

Error contract:
- Exit 0 on success, 2 on unknown node/field, 3 on malformed input.
- BdtInputError (exit 3) and BdtNotFoundError (exit 2) classes
  centralize error handling so the CLI shell stays thin.
- Field-trace narrowing refinements prevent false positives from
  substring matches in formula bodies.
- Upstream/downstream walkers harden against broken refs discovered
  during internal-BDT audit.

Design invariants:
- Python owns truth (parsing, DAG math, field discovery). LLM owns
  narrative (explaining what the structure means to a user).
- No external dependencies — stdlib only: argparse, json, pathlib,
  re, sys, collections, dataclasses, typing.
- Output size budgets: every subcommand caps its default-mode output
  so summaries fit in a single LLM context window; --json dumps
  everything for agents that need raw data.

@W-22196528@
2026-04-23 23:29:05 +05:30