afv-library/skills/explaining-batch-data-transform/references/bdt-function-catalog.md

125 lines
5.7 KiB
Markdown
Raw Normal View History

feat(bdt): reference docs + sample BDTs @W-22196528@ Adds the curated reference library the skill loads on demand, plus four synthetic sample BDTs used by docs, tests, and LLM-mode demos. references/ (4 curated Markdown files): - bdt-reference.md — top-level BDT JSON anatomy: envelope, nodes, edges, UI layer, definitions, businessType semantics. Cites the core-262 upstream JSON schema and Connect API spec. - bdt-node-catalog.md — every node type (DMO Source, DMO Sink, Filter, Join, Union, Aggregate, Window, Formula, Split, Append, etc.) with its required/optional fields and typical usage. Audited against core-262 enums. - bdt-function-catalog.md — the expression-language function surface (string, numeric, date, conditional, aggregate). Grouped by category with signature + one-line semantics. - bdt-window-functions.md — windowing operators (ROW_NUMBER, RANK, LEAD/LAG, running aggregates) with PARTITION BY / ORDER BY grammar and gotchas. assets/sample_bdts/ (4 synthetic, dependency-free BDTs): - minimal_dmo_to_dmo.json — smallest valid BDT (1 source, 1 sink). - joins_and_filters.json — join + filter composition. - window_and_aggregate.json — window function + aggregate in one graph. - append_and_split.json — append-then-split branching topology. Grounding rules enforced in this commit: - Every claim in references/ cites an upstream source (core-262 JSON schema, Connect API reference, or the Data Cloud BDT editor spec). No speculative content. - No raw DITA or internal-only documentation is shipped; references are synthesized from public-facing material. - BusinessTypeEnum values use the canonical camelCase casing from core-262 (case-cleanup fix included here). - Sample BDTs are original synthetic fixtures, not redacted customer data. Each is small enough to read end-to-end. @W-22196528@
2026-04-24 01:15:27 +08:00
# SFSQL Function Catalog — for `formula` reasoning
> **Last synced:** 2026-04-23 from the Data Processing Engine PDF (2025-06-24 version) and
> the BDT help DITA XMLs (`c360_a_batch_transform_numeric.xml`, `_string.xml`,
> `_date_functions.xml`, `_boolean_functions.xml`, `_multivalue_functions.xml`,
> `_additionalfunctions.xml`). Consult this file before interpreting any `formulaExpression`.
## How formulas appear in BDT JSON
```jsonc
"action": "formula",
"parameters": {
"expressionType": "SQL", // or "DCSQL"
"fields": [
{
"name": "OutputField__c",
"label": "Output Label",
"type": "TEXT", // TEXT | NUMBER | BOOLEAN | DATE_ONLY | DATETIME
"businessType": "Text", // user-visible type name
"formulaExpression": "upper(coalesce(SourceField__c, ''))",
"precision": 255,
"scale": 0, // only for NUMBER
"defaultValue": ""
}
]
}
```
## Functions by family
### Math / Numeric
| Function | Signature | Purpose |
|---|---|---|
| `abs(n)` | number → number | Absolute value (strips sign). |
| `ceiling(n)` | number → number | Round up, away from zero for negatives. |
fix: address PRizm round-2 review findings from internal port @W-22196528@ Ports the 14 applicable round-2 fixes from the internal PRizm review to the external afv-library source. (plugin.json finding #13 is internal-only and does not apply here.) Accepted fixes: 1 bdt_analyze.py select_definition: isinstance(index, int) -> type(index) is int (reject bool subclass of int). 2 bdt_analyze.py cmd_formula: precompute upstream fields_produced map to eliminate O(consumed x upstream) recomputation. 3 bdt_analyze.py topo_order: replace sort-on-every-iteration with deque-based Kahn (preserves deterministic order). 4 test_bdt_analyze.py test_all_subcommands_work_on_api_input: wrap loop in self.subTest for per-iteration failure reporting. 5 test_bdt_analyze.py test_invalid_json_raises: register addCleanup BEFORE write_text so cleanup runs even if write fails. 6 SKILL.md: remove misleading "strip __c suffix" field-trace troubleshooting advice; replace with source-DMO passthrough guidance. 7 SKILL.md: rewrite cycle error message to direct user to Data Cloud viewer (re-export will not fix a cycle in the BDT definition). 8 SKILL.md: add explicit --definition N flag documentation and list every subcommand that accepts it. 9 SKILL.md: refine I2 routing guidance — pick the earliest formula/ computeRelative node, not the output mapping. 11 bdt-function-catalog.md: clean floor() wording ("always rounds down on the number line"). 12 bdt-node-catalog.md: businessType enum — assert list is canonical/complete; instruct skill to surface+flag any unknown value. 14 append_and_split.json: change split subject from fragile CustomerFullName__c split-on-space to ChannelOrderKey__c split-on-hyphen; avoids name-parsing pitfall (middle names / multi-space). 15 window_and_aggregate.json: use the computed OrderRank__c via a new FIRST_ORDER_AMOUNT formula node (case when rank=1 then amount else 0) feeding AGG_BY_ACCOUNT; illustrates the rank+formula+aggregate idiom. 16 joins_and_filters.json: add explicit IS_NOT_NULL filter expression on GrandTotalAmount__c (defense-in-depth beyond GREATER_THAN 0). Rejected finding (rationale posted as PR comment): 10 bdt-node-catalog.md split section — reviewer claimed split is a pipeline-branching node. Canonical sources (core-262 SplitParametersInputRepresentation.java + Salesforce help DITA c360_a_batch_transform_split.xml) both describe split as a string-splitting operation. Current docs are correct; not changing. Tests: 92/92 passing.
2026-04-24 13:25:18 +08:00
| `floor(n)` | number → number | Round toward negative infinity. Always rounds down on the number line: `floor(2.7) = 2`, `floor(-2.3) = -3`. |
feat(bdt): reference docs + sample BDTs @W-22196528@ Adds the curated reference library the skill loads on demand, plus four synthetic sample BDTs used by docs, tests, and LLM-mode demos. references/ (4 curated Markdown files): - bdt-reference.md — top-level BDT JSON anatomy: envelope, nodes, edges, UI layer, definitions, businessType semantics. Cites the core-262 upstream JSON schema and Connect API spec. - bdt-node-catalog.md — every node type (DMO Source, DMO Sink, Filter, Join, Union, Aggregate, Window, Formula, Split, Append, etc.) with its required/optional fields and typical usage. Audited against core-262 enums. - bdt-function-catalog.md — the expression-language function surface (string, numeric, date, conditional, aggregate). Grouped by category with signature + one-line semantics. - bdt-window-functions.md — windowing operators (ROW_NUMBER, RANK, LEAD/LAG, running aggregates) with PARTITION BY / ORDER BY grammar and gotchas. assets/sample_bdts/ (4 synthetic, dependency-free BDTs): - minimal_dmo_to_dmo.json — smallest valid BDT (1 source, 1 sink). - joins_and_filters.json — join + filter composition. - window_and_aggregate.json — window function + aggregate in one graph. - append_and_split.json — append-then-split branching topology. Grounding rules enforced in this commit: - Every claim in references/ cites an upstream source (core-262 JSON schema, Connect API reference, or the Data Cloud BDT editor spec). No speculative content. - No raw DITA or internal-only documentation is shipped; references are synthesized from public-facing material. - BusinessTypeEnum values use the canonical camelCase casing from core-262 (case-cleanup fix included here). - Sample BDTs are original synthetic fixtures, not redacted customer data. Each is small enough to read end-to-end. @W-22196528@
2026-04-24 01:15:27 +08:00
| `exp(n)` | number → number | e raised to n. |
| `log(base, n)` | number → number | Logarithm of n in the given base. |
| `max(a, b, …)` | number → number | Largest value. |
| `min(a, b, …)` | number → number | Smallest value. |
| `mod(a, b)` | number → number | Remainder after dividing a by b. |
| `power(a, b)` | number → number | a raised to the power b. |
| `round(n, digits)` | number → number | Round to `digits` places. |
| `sqrt(n)` | number → number | Positive square root. |
| `trunc(n, digits)` | number → number | Truncate to `digits` places (doesn't round). |
### String
| Function | Signature | Purpose |
|---|---|---|
| `begins(s, prefix)` | text → boolean | True if s starts with prefix. |
| `concat(a, b, …)` | text → text | Concatenate. |
| `contains(haystack, needle)` | text → boolean | True if haystack contains needle. |
| `ends(s, suffix)` | text → boolean | True if s ends with suffix. |
| `length(s)` | text → number | Character count. |
| `lower(s)` | text → text | Lowercase (locale-aware if a locale is provided). |
| `ltrim(s)` / `ltrim(s, substring)` | text → text | Remove leading whitespace (or a specific substring). |
| `rtrim(s)` / `rtrim(s, substring)` | text → text | Remove trailing whitespace (or substring). |
| `substitute(s, old, new)` | text → text | Replace old with new in s. |
| `substr(s, start, length)` | text → text | Extract a substring. |
| `text(n)` | number → text | Convert a number to its text form. |
| `trim(s)` / `trim(s, substring)` | text → text | Remove leading + trailing whitespace or substring. |
| `upper(s)` | text → text | Uppercase. |
| `uuid()` | → text | Newly generated unique ID. |
| `value(s)` | text → number | Parse a text representation of a number into a numeric value. |
### Date
| Function | Signature | Purpose |
|---|---|---|
| `adddays(date, n)` | date → date | Add n days. |
| `addmonths(date, n)` | date → date | Add n months. |
| `datediff(start, end)` | date × date → number | Days between two dates. |
| `datetimevalue(text_or_date)` | → datetime | GMT/UTC date+time value. |
| `datevalue(text_or_datetime)` | → date | Extract the date part. |
| `day(date)` | date → number | 1-31. |
| `monthdiff(start, end)` | date × date → number | Months between two dates. |
| `now()` | → datetime | Current moment (UTC). |
| `today()` | → date | Current date. |
| `weekday(date)` | date → number | 1=Sunday, 2=Monday, … 7=Saturday. |
### Boolean / Logical
| Function | Signature | Purpose |
|---|---|---|
| `and(a, b, …)` | bool → bool | True when all are true. |
| `or(a, b, …)` | bool → bool | True when any is true. |
| `if(cond, then_val, else_val)` | → any | Ternary. |
| `case when <cond1> then <val1> else <elseval> end` | — | Multi-branch (SQL CASE). |
| `blankvalue(expr, substitute)` | → any | Substitute when expr is blank. |
| `isblank(expr)` | → bool | True when expr is blank. |
| `isnull(expr)` | → bool | True when expr is null. |
| `nullvalue(expr, substitute)` | → any | Substitute when expr is null; returns expr otherwise. |
### Multivalue
| Function | Signature | Purpose |
|---|---|---|
| `sequence(start, end, step?)` | → array | Array of numbers/dates between start and end (step defaults to 1 / 1 day). |
| `explode(array)` | → row per element | Convert multivalue data into one row per element. **Cannot be nested inside another function.** |
### Additional / Null-handling (commonly seen)
| Function | Signature | Purpose |
|---|---|---|
| `coalesce(a, b, …)` | → any | First non-null argument (SQL standard). |
## Patterns to recognize in narration
- **`coalesce(X, Y)`** — "use X when present, otherwise Y." Very common for default-value
handling.
- **`case when RANK = 1 then X else 0 end`** — the "first-occurrence extract" idiom paired
with a `computeRelative` `row_number()` node.
- **`case when FLAG = 'Y' then 'Y' else 'N' end`** — boolean-like text normalization.
- **`concat(coalesce(A, ''), coalesce(B, ''))`** — safe string concatenation that defaults
NULLs to empty strings (producing a synthetic composite key).
## Sources
- Data Processing Engine reference PDF (captured at `research/bdt-doc-10-data_processing_engine_6-24-2025.pdf.md`)
- BDT help XML: `c360_a_batch_transform_numeric.xml`, `_string.xml`, `_date_functions.xml`,
`_boolean_functions.xml`, `_multivalue_functions.xml`, `_additionalfunctions.xml` (captured
at `research/help-xml/`).