afv-library/skills/investigating-agentforce-d360/assets/dc/gateway_requests.sql
rjayagopal 20ae436442 @W-22707610 feat: add investigating-agentforce-d360 skill
Data Cloud 360° view of a single Agentforce session — DC-only, zero
Splunk dependency. Pulls 24 STDM + GenAI DMOs via the Data Cloud Query
REST API, assembles a hierarchical session tree (Interaction → Step →
Generation → GatewayRequest), and renders a human-readable markdown
summary with transcript + per-turn topic/action invocations + LLM
generations + tool calls + audit chain.

Migrated as a standalone Apache-2.0 skill from an internal hub plugin —
self-contained, no sibling-skill or plugin dependencies.

What this skill answers:
  - "Trace session <uuid>" / "Summarize what happened in <0Mw…>"
  - "Find escalated sessions today on Messaging in <org>"
  - Session discovery by time / agent / channel / outcome / conversation
    text when the user has no session id

What it does NOT answer (use a different surface):
  - Design-time architecture — use investigating-agentforce-architecture
  - Runtime planner availability — DC alone can't tell you which
    topic/action was eligible for the classifier on a given turn

Skill layout:
  - 8 Python pipeline modules (fetch_dc, assemble_dc, render_dc,
    discover_sessions, resolve_session, dc, storage, config)
  - 4 _shared helpers (paths, fs_guard, sql, __init__) with skill-scoped
    DATA_ROOT (~/.claude/data/investigating-agentforce-d360/)
  - 26 SQL templates under assets/dc/
  - 27 test files (367 tests + 18 subtests, 100% passing)
  - 3 reference docs (artifacts.md, dc_dmo_fields.md,
    dc_pipeline_contract.md)
  - SKILL.md (sf-skills frontmatter, license: Apache-2.0,
    metadata.version: "1.0")
  - README.md (external-facing quick-start)
  - tools/grant_allowlist.py (idempotent first-run permission grant)
  - tools/archive_data_dir.sh (opt-in stop-hook tarballer)

Quality gates:
  - pytest scripts/tests/: 367 passed + 18 subtests, 0 failures
  - npm run validate:skills: 62 of 62 skill(s) checked, 0 errors
  - Live end-to-end runs against 3 real Salesforce sessions exercising
    both the full-tree and STDM-lag gateway-direct render branches
  - 4 independent code-review rounds (correctness, security, markdown,
    architecture-critic) — all findings addressed

Customer-data hygiene: no live tenant ids, no internal sprint markers,
no hub/sibling-skill references. Synthetic fixtures look obviously
synthetic (`019dface-…` UUIDs, `0MwTESTMSG…` MessagingSession ids,
`00DTESTORG…` org ids, `MyAgent` placeholder agent name).

Sibling skill: investigating-agentforce-architecture (PR #278) — same
migration pattern, design-time metadata; complementary scope.
2026-05-28 20:57:52 +10:00

90 lines
3.4 KiB
SQL

-- Gateway requests — one row per LLM request at the GenAI Gateway.
-- DMO: GenAIGatewayRequest__dlm
--
-- Placeholders (substituted by scripts/dc.py.load_sql):
-- WHERE_CLAUSE — the filter expression, no "WHERE" keyword
-- ORDER_BY — full "ORDER BY <col>" or empty string
--
-- Richer than GenAIGeneration — carries the actual prompt text, token counts,
-- model params (temperature/penalties), session/user IDs, bot version, and
-- the masked-prompt variant.
--
-- NOTE: No `ssot__` prefix — fields end in `__c` directly.
--
-- sessionId__c storage format (verified live): the value is stored as a
-- literal 40-char string INCLUDING surrounding double-quotes, e.g.
-- sessionId__c = "<session_uuid>"
-- Non-session features (prompt-builder previews, eval harnesses, etc.) store
-- the literal sentinel "no_session". Exact-match queries MUST include the
-- double-quotes:
-- WHERE sessionId__c = '"<session_uuid>"'
-- Or use LIKE with wildcards (robust against format variants):
-- WHERE sessionId__c LIKE '%<session_uuid>%'
-- Raw-UUID exact match returns 0 rows — the quotes are part of the stored value.
--
-- Forward join path from a session:
-- Session.ssot__Id__c → GatewayRequest.sessionId__c (LIKE or quoted match)
-- This is the authoritative and only supported entry point. GatewayRequest is
-- then the parent for all downstream audit-chain children — Response (via
-- generationRequestId__c), Tag/ObjRecord/Metadata/LLM (via parent__c).
-- See `scripts/fetch_dc.py` wave 3 and `references/dc_dmo_fields.md` "Cross-DMO
-- join map" for the full forward tree.
SELECT
gatewayRequestId__c,
generationGroupId__c,
sessionId__c,
userId__c,
botVersionId__c,
plannerId__c,
feature__c,
appType__c,
model__c,
provider__c,
promptTemplateDevName__c,
promptTemplateVersionNo__c,
prompt__c,
maskedPrompt__c,
parameters__c,
temperature__c,
frequencyPenalty__c,
presencePenalty__c,
stopSequences__c,
numGenerations__c,
promptTokens__c,
completionTokens__c,
totalTokens__c,
enableInputSafetyScoring__c,
enableOutputSafetyScoring__c,
enablePiiMasking__c,
timestamp__c,
orgId__c,
cloud__c
FROM GenAIGatewayRequest__dlm
WHERE {{WHERE_CLAUSE}}
{{ORDER_BY}};
-- ============================================================================
-- EXAMPLE WHERE clauses (pass via where_clause=)
-- ============================================================================
-- Requests for one session (direct FK; note the mandatory double-quoted form).
-- Two equivalent WHERE forms — both verified live, both return the same rows:
-- WHERE → sessionId__c = '"<session_id>"' (exact match on quoted string)
-- WHERE → sessionId__c LIKE '%<session_id>%' (format-tolerant)
-- ORDER BY → ORDER BY timestamp__c
-- Requests for a specific set of gatewayRequestIds (e.g. narrowing after a
-- session fetch, or lookup by ids harvested from another query)
-- WHERE → gatewayRequestId__c IN ('<req_id1>','<req_id2>',...)
-- All requests for a bot version in a time window
-- WHERE → botVersionId__c = '<version_id>'
-- AND timestamp__c >= '<iso_cutoff>'
-- ORDER BY → ORDER BY timestamp__c DESC
-- Requests using a specific prompt template
-- WHERE → promptTemplateDevName__c = '<template_dev_name>'
-- AND timestamp__c >= '<iso_cutoff>'