mirror of
https://github.com/forcedotcom/afv-library.git
synced 2026-08-09 00:42:46 +08:00
67d919947f
2 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
67d919947f
|
fix: address PRizm round-2 review findings from internal port @W-22196528@
Ports the 14 applicable round-2 fixes from the internal PRizm review to the external afv-library source. (plugin.json finding #13 is internal-only and does not apply here.) Accepted fixes: 1 bdt_analyze.py select_definition: isinstance(index, int) -> type(index) is int (reject bool subclass of int). 2 bdt_analyze.py cmd_formula: precompute upstream fields_produced map to eliminate O(consumed x upstream) recomputation. 3 bdt_analyze.py topo_order: replace sort-on-every-iteration with deque-based Kahn (preserves deterministic order). 4 test_bdt_analyze.py test_all_subcommands_work_on_api_input: wrap loop in self.subTest for per-iteration failure reporting. 5 test_bdt_analyze.py test_invalid_json_raises: register addCleanup BEFORE write_text so cleanup runs even if write fails. 6 SKILL.md: remove misleading "strip __c suffix" field-trace troubleshooting advice; replace with source-DMO passthrough guidance. 7 SKILL.md: rewrite cycle error message to direct user to Data Cloud viewer (re-export will not fix a cycle in the BDT definition). 8 SKILL.md: add explicit --definition N flag documentation and list every subcommand that accepts it. 9 SKILL.md: refine I2 routing guidance — pick the earliest formula/ computeRelative node, not the output mapping. 11 bdt-function-catalog.md: clean floor() wording ("always rounds down on the number line"). 12 bdt-node-catalog.md: businessType enum — assert list is canonical/complete; instruct skill to surface+flag any unknown value. 14 append_and_split.json: change split subject from fragile CustomerFullName__c split-on-space to ChannelOrderKey__c split-on-hyphen; avoids name-parsing pitfall (middle names / multi-space). 15 window_and_aggregate.json: use the computed OrderRank__c via a new FIRST_ORDER_AMOUNT formula node (case when rank=1 then amount else 0) feeding AGG_BY_ACCOUNT; illustrates the rank+formula+aggregate idiom. 16 joins_and_filters.json: add explicit IS_NOT_NULL filter expression on GrandTotalAmount__c (defense-in-depth beyond GREATER_THAN 0). Rejected finding (rationale posted as PR comment): 10 bdt-node-catalog.md split section — reviewer claimed split is a pipeline-branching node. Canonical sources (core-262 SplitParametersInputRepresentation.java + Salesforce help DITA c360_a_batch_transform_split.xml) both describe split as a string-splitting operation. Current docs are correct; not changing. Tests: 92/92 passing. |
||
|
|
9f1603ed2f
|
feat(bdt): Python parser core + CLI subcommands @W-22196528@
Adds scripts/bdt_analyze.py — a generic, stdlib-only DAG parser and
query CLI for Salesforce Data Cloud BDT JSON. 1,379 lines.
Parser primitives:
- DataTransform.from_path / from_dict — accepts three input shapes:
editor export ({version, nodes, ui, ...}), Connect API
single-definition ({name, label, type, definition: {...}}), and
Connect API multi-definition ({name, definitions: [{name, label,
type, definition}, ...]}).
- Node dataclass with ui_label / ui_description fallback resolution.
- roots() / sinks() / topo_order() — Kahn's algorithm with
deterministic tie-break for reproducible output.
- upstream() / downstream() traversal resilient to broken refs and
cycles (does not infinite-loop on self-edges or cycles).
- broken_references(), fields_produced(), fields_consumed(),
_scrape_field_refs() heuristics for field-trace discovery.
CLI (argparse, 10 subcommands, each supports --json for machine
output):
- summary — node counts, source/sink counts, stage totals.
- sources — list source nodes (no upstream).
- outputs — list sink nodes (no downstream).
- stages — topologically ordered stages.
- nodes — flat node listing with labels.
- node <name> — per-node detail (action, inputs, outputs, fields,
UI label/description).
- lineage <node> — upstream + downstream chain from a node.
- field-trace <field> — which nodes produce/consume a given field.
- formula <node> — extract formulas/expressions from a node.
- definitions — lists definitions in multi-definition
payloads; every other subcommand accepts
--definition N (default 0) to route into a
specific definition within the payload.
Error contract:
- Exit 0 on success, 2 on unknown node/field, 3 on malformed input.
- BdtInputError (exit 3) and BdtNotFoundError (exit 2) classes
centralize error handling so the CLI shell stays thin.
- Field-trace narrowing refinements prevent false positives from
substring matches in formula bodies.
- Upstream/downstream walkers harden against broken refs discovered
during internal-BDT audit.
Design invariants:
- Python owns truth (parsing, DAG math, field discovery). LLM owns
narrative (explaining what the structure means to a user).
- No external dependencies — stdlib only: argparse, json, pathlib,
re, sys, collections, dataclasses, typing.
- Output size budgets: every subcommand caps its default-mode output
so summaries fit in a single LLM context window; --json dumps
everything for agents that need raw data.
@W-22196528@
|