afv-library/skills/explaining-batch-data-transform/assets/sample_bdts/append_and_split.json
Gaurav Bajpai db1aca7221
fix: address PRizm review findings from internal port PR @W-22196528@
Addresses 4 critical findings surfaced during PRizm code review of the
internal plugin port of this skill (internal PR #19). Applying the same
fixes here keeps the external canonical source and the internal port in
sync.

1. Test cleanup discipline: switch `test_invalid_json_raises` from
   `try/finally` to `self.addCleanup(p.unlink, missing_ok=True)` — the
   unittest-idiomatic way to guarantee temp-file cleanup regardless of
   how the test exits.

2. `floor()` description in bdt-function-catalog.md: the old row was
   self-contradictory ("toward zero" AND "toward next integer up" in
   the same cell). Replace with a single coherent definition: rounds
   toward negative infinity; for negatives rounds away from zero
   (e.g., `floor(-2.3) = -3`).

3. Split-node documentation in bdt-node-catalog.md: the old doc claimed
   `split` routes rows into downstream branches via `branches[]` with
   per-branch predicates. That is not the canonical schema. Per
   `SplitParametersInputRepresentation` in core-262-public, `split` is
   a string-splitting operation: one `sourceField` + `delimiter` →
   N `targetFields` (one row in, one row out; columns added). Rewrote
   the section with the correct parameters, lineage effect, gotchas,
   and a canonical example. Row-routing belongs in `filter` nodes.

4. Sample `assets/sample_bdts/append_and_split.json`: the old sample
   used the invented `branches[]` shape AND routed the same split into
   two downstream outputs that each expected different rows — which is
   not how `split` works. Rewrote the sample so:
     - `appendV2` unions two order sources (unchanged intent).
     - `split` uses canonical `{sourceField, delimiter, targetFields}`
       splitting `CustomerFullName__c` into first + last name columns.
     - One downstream output consumes the new columns (removes the
       fake two-branch fan-out).

Tests: 92/92 passing. Sample parses and runs through `bdt_analyze.py
summary` cleanly (5 nodes: 2 load + 1 appendV2 + 1 split + 1 outputD360).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-24 10:19:36 +05:30

80 lines
3.5 KiB
JSON

{
"version": "66.0",
"nodes": {
"LOAD_WEB_ORDERS": {
"action": "load",
"sources": [],
"parameters": {
"dataset": {"name": "WebOrders__dlo", "type": "dataLakeObject"},
"fields": ["OrderId__c", "Amount__c", "Channel__c", "CustomerFullName__c"],
"sampleDetails": {"type": "TopN", "sortBy": []}
}
},
"LOAD_STORE_ORDERS": {
"action": "load",
"sources": [],
"parameters": {
"dataset": {"name": "StoreOrders__dlo", "type": "dataLakeObject"},
"fields": ["OrderId__c", "Amount__c", "Channel__c", "CustomerFullName__c"],
"sampleDetails": {"type": "TopN", "sortBy": []}
}
},
"APPEND_ALL_ORDERS": {
"action": "appendV2",
"sources": ["LOAD_WEB_ORDERS", "LOAD_STORE_ORDERS"],
"parameters": {
"allowImplicitDisjointSchema": false,
"fieldMappings": [
{"targetField": "OrderId__c", "sources": [{"node": "LOAD_WEB_ORDERS", "field": "OrderId__c"}, {"node": "LOAD_STORE_ORDERS", "field": "OrderId__c"}]},
{"targetField": "Amount__c", "sources": [{"node": "LOAD_WEB_ORDERS", "field": "Amount__c"}, {"node": "LOAD_STORE_ORDERS", "field": "Amount__c"}]},
{"targetField": "Channel__c", "sources": [{"node": "LOAD_WEB_ORDERS", "field": "Channel__c"}, {"node": "LOAD_STORE_ORDERS", "field": "Channel__c"}]},
{"targetField": "CustomerFullName__c", "sources": [{"node": "LOAD_WEB_ORDERS", "field": "CustomerFullName__c"}, {"node": "LOAD_STORE_ORDERS", "field": "CustomerFullName__c"}]}
]
}
},
"SPLIT_CUSTOMER_NAME": {
"action": "split",
"sources": ["APPEND_ALL_ORDERS"],
"parameters": {
"sourceField": "CustomerFullName__c",
"delimiter": " ",
"targetFields": [
{"name": "CustomerFirstName__c", "label": "Customer First Name"},
{"name": "CustomerLastName__c", "label": "Customer Last Name"}
]
}
},
"OUTPUT_ORDERS": {
"action": "outputD360",
"sources": ["SPLIT_CUSTOMER_NAME"],
"parameters": {
"name": "OrdersWithCustomerName__dlm",
"type": "dataModelObject",
"writeMode": "OVERWRITE",
"fieldsMappings": [
{"sourceField": "OrderId__c", "targetField": "OrderId__c"},
{"sourceField": "Amount__c", "targetField": "Amount__c"},
{"sourceField": "Channel__c", "targetField": "Channel__c"},
{"sourceField": "CustomerFirstName__c", "targetField": "FirstName__c"},
{"sourceField": "CustomerLastName__c", "targetField": "LastName__c"}
]
}
}
},
"ui": {
"nodes": {
"LOAD_WEB_ORDERS": {"label": "Web Orders", "type": "LOAD_DATASET", "top": 100, "left": 100},
"LOAD_STORE_ORDERS": {"label": "Store Orders", "type": "LOAD_DATASET", "top": 260, "left": 100},
"APPEND_ALL_ORDERS": {"label": "Union", "type": "APPEND", "top": 180, "left": 260},
"SPLIT_CUSTOMER_NAME": {"label": "Split full name", "type": "SPLIT", "top": 180, "left": 420},
"OUTPUT_ORDERS": {"label": "Orders with name", "type": "OUTPUT", "top": 180, "left": 580}
},
"connectors": [
{"source": "LOAD_WEB_ORDERS", "target": "APPEND_ALL_ORDERS"},
{"source": "LOAD_STORE_ORDERS", "target": "APPEND_ALL_ORDERS"},
{"source": "APPEND_ALL_ORDERS", "target": "SPLIT_CUSTOMER_NAME"},
{"source": "SPLIT_CUSTOMER_NAME", "target": "OUTPUT_ORDERS"}
]
}
}