Repository navigation
feat(sensor): add opt-in GenAI semconv attributes to OTel session logs - #161
Open
Zhuoxi2000 wants to merge 2 commits into
Open
Zhuoxi2000 wants to merge 2 commits into
Zhuoxi2000 wants to merge 2 commits into
Conversation
Add a gen_ai_attributes OpenTelemetry config option. When enabled, adr.agent.session log records also carry gen_ai.conversation.id and gen_ai.agent.name so GenAI observability backends can correlate ADR sessions without a Collector transform. The body and adr.* attributes are unchanged, and the option is off by default. The default value is left out of the delivery checkpoint fingerprint so existing checkpoints stay valid after upgrade.
gen_ai.conversation.id now carries each harness's own session ID instead of the prefixed ADR session ID, so ADR sessions join with the telemetry the harness reports itself. adr.session.id and the body are unchanged. The exporter maps each source back from the prefix its parser adds. Claude subagents get their parent's sessionId, Claude Desktop uses cliSessionId, and opencode gets its "ses_" prefix back. The attribute is omitted rather than guessed when the source is unknown, the prefix does not match, or the native ID is unavailable. The README documents the mapping per source. Tests run each parser on a fixture and check the exported ID, and fail if a new parser has no mapping entry.
|
|
barisozbas
approved these changes
Oct 6, 2026
Collaborator
There was a problem hiding this comment.
LGTM!
@Zhuoxi2000 Please sign the CLA so that it can be merge.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What type of PR is this? (check all applicable)
Related issue: Closes #153
What changed?
exporters/config.py: new optional keygen_ai_attributes. It is a strict boolean and defaults tofalse.exporters/opentelemetry.py: when the key is on,adr.agent.sessionrecords also carry:gen_ai.agent.name: the source, the same value asadr.source.gen_ai.conversation.id: the harness's own session ID.The body, all
adr.*attributes (includingadr.session.id), system-configuration records and health records are unchanged.How the ID is mapped: every parser builds
session_idas<source prefix><native id>. A table keyed byentry.source(_CONVERSATION_ID_BY_SOURCE) undoes that prefix. Mapping in the exporter meansAgentEventgets no new field, so payloads and per-session checkpoint hashes stay the same. The attribute is left out when:session_iddoesn't have the expected prefix;It never falls back to
adr.session.id.exporters/delivery_checkpoint.py: the new config field is left out of the destination fingerprint while it isfalse, so existing.adr-otel-delivery.<hash>.jsoncheckpoints stay valid after upgrade. Turning it on changes the destination and resends once.README.md: documents the option, a per-source mapping table and the omission rule.session_idgen_ai.conversation.idantigravity_<id><id>: conversation directory nameclaude_<id>/claude_<id>_agent_<agent><id>: JSONLsessionId. Subagents get their parent'ssessionId, matching Claude Code's own telemetryclaude_desktop_[dispatch_]<uuid>cliSessionId, omitted when absentcline_<id><id>: task directory namecodex_<id><id>:session_meta.payload.idcopilot_<id><id>:session.start/session.resumesessionId, falling back to the session-state directory name, as the parser doescursor_<id><id>: composer IDdsh_<id><id>: session headeridgemini_<id><id>:sessionId. Subagents keep their ownopencode_<id>ses_<id>warp_<id><id>:agent_conversations.conversation_idThree rows involve a judgment call, and I'm happy to change any of them:
sessionIdprefix (local_/local_ditto_). This PR usescliSessionIdinstead. It is kept verbatim insession_context.cli_session_id, and it is the ID that matches Claude Code CLI telemetry. Sessions without it get nogen_ai.conversation.id. The alternative is to rebuildlocal_[ditto_]<uuid>from the ADR ID.ses_only when it is present. Its docstring says opencode session IDs always start withses_, so this PR adds it back. If you'd rather not rely on that, the alternative is to omit the attribute for opencode.session_context.conversation_id.Why?
As discussed in #153, GenAI and agent-observability backends correlate on the OpenTelemetry GenAI registry attributes, and a Collector shouldn't be required just for these.
gen_ai.conversation.idcarries the harness-native ID, so ADR sessions can join the session IDs that harnesses report in their own telemetry. Model, provider and token usage are left out for the reasons in the issue.How did you test it?
New
tests/test_gen_ai_mapping.py:gen_ai_attributeson. It then asserts thatgen_ai.conversation.idequals the raw native ID in the fixture, that the body equalsget_non_null_fields(), and thatadr.session.idis unchanged. If a parser later changes its prefix, these tests fail.cliSessionId, the Copilot directory fallback, Gemini subagents, and opencodeses_.test_unavailable_conversation_id_is_omitted: unknown source, wrong prefix, empty native ID, and a Claude subagent without a parent session. Each yields nogen_ai.conversation.id, whilegen_ai.agent.nameand the exactadr.*set remain. Claude Desktop withoutcliSessionIdis covered separately, through the parser.test_every_parser_has_a_conversation_id_mapping: every parser exported fromadr_sensor.parsersneeds an entry in the mapping table, so adding a harness forces a decision.Other new tests:
test_export_adds_gen_ai_attributes_when_enabled[gpt-test|None]: the exact attribute set when the option is on.test_load_opentelemetry_config_defaults/_all_fields(extended) and two new non-boolean cases intest_load_opentelemetry_config_rejects_invalid_values.test_enabling_gen_ai_attributes_resends_sessions.test_default_gen_ai_attributes_keep_existing_checkpoint: pins the checkpoint filename thatmainproduces for the default config.test_export_preserves_complete_agent_event_payloadis unchanged and still pins the default attribute set and body.With these tests on top of
main's source,test_gen_ai_mapping.pyfails at import, and all the other new or extended tests except the checkpoint-name one fail (7 failures, for exampleunexpected keyword argument 'gen_ai_attributes'andunknown OpenTelemetry config field(s): gen_ai_attributes). The remaining 57 tests intest_opentelemetry_exporter.py,test_exporter_config.pyandtest_delivery_checkpoint.pypass. They includetest_default_gen_ai_attributes_keep_existing_checkpoint, which shows that the pinned checkpoint name is the onemainproduces.Against the first commit in this PR, where
gen_ai.conversation.idwas still the ADR session ID,test_gen_ai_mapping.pyalso fails at import, because_CONVERSATION_ID_BY_SOURCEdoesn't exist yet. With that import stubbed so the module loads, 21 of the 42 tests intest_gen_ai_mapping.pyandtest_opentelemetry_exporter.pyfail, for exampleassert 'codex_sess1' == 'sess1'.Potential risks
None while the option is off: attributes, body and checkpoint paths are identical to
main. When it is on:session_idprefix. The per-harness tests catch a prefix change.