Skip to content

feat: enable default context compression in SDK and Studio - #1152

Open
zyn080302 wants to merge 20 commits into
volcengine:mainfrom
zyn080302:feat/default-context-compression
Open

zyn080302 wants to merge 20 commits into
volcengine:mainfrom
zyn080302:feat/default-context-compression

Conversation

@zyn080302

@zyn080302 zyn080302 commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

This change enables recoverable context compression by default in the veADK SDK and Studio, reducing input-limit errors caused by long conversations and large tool responses. Model requests retain relevant evidence and source references, while complete session records remain stored in SQLite by default.

Agents can retrieve original content when needed without rerunning business tools. Compression thresholds are configurable, and developers can explicitly disable compression. Bounded retrieval, local fallback, and final request-size checks improve reliability.

The branch includes upstream main 4d37da98, preserving the Studio A2A registry deployment validation fixes and updated default UI text. Packaged Studio assets match the combined frontend source. Documentation, examples, README and tutorial updates are outside this PR; the SDK scenario formerly exercised through an example remains covered directly in the Runner integration suite.

The A2A deployment regression deterministically exercises asynchronous capability preparation (202), polls with a deadline, and still requires a successful terminal capability and deployment result. This removes an immediate-completion assumption that failed under concurrent CI load.

Validation:

  • Isolated context compression gate after scope reduction: 1,448 passed, 4 skipped (Python 3.12 / google-adk 2.2.0); all assertions from the migrated example regression retained.
  • A2A deployment and release-gate contracts: 11 passed. The original 202 == 200 CI failure was reproduced locally before the test correction.
  • Ruff check/format and staged secret scan passed.
  • Unchanged frontend/bundle validation from the preceding integration commit: 1,246 frontend tests, 22 Sidecar coverage tests, TypeScript and both production builds passed; 113 packaged files and 350 internal references verified; both locales and all 21 i18n namespaces consistent.

@zyn080302
zyn080302 force-pushed the feat/default-context-compression branch from 465b862 to db69ba3 Compare September 28, 2026 12:11
@zyn080302
zyn080302 force-pushed the feat/default-context-compression branch from 9cfee74 to 9cf4d93 Compare October 9, 2026 09:12
yanan.zhangyn added 19 commits October 9, 2026 17:13
…tudio

Preserve original sessions in project SQLite and select budgeted evidence for
standard Agents without requiring manual retriever setup. Expose input limits
and compression thresholds consistently in both Studio creation flows and
in generated Agent configuration. Bound optional embedding work, preserve
provider credential boundaries, and retain source lookup after restart.

Validation on authorized Devbox: SDK contracts 1157 passed / 5 skipped;
frontend 1245 passed; production app/widget build and asset references pass;
Ruff and pre-commit including gitleaks pass for source and bundled WebUI.
Core context module type audit has no diagnostics; broad Pyright still
reports upstream and test-double diagnostics. No live model evaluation or
customer Runtime changes. Upstream base: 763865d.
Bound base and eval dependencies to google-adk >=1.34,<2.3 while AgentKit 0.8.x requires OpenTelemetry <=1.37. Keep the full gate for supported targets and add regressions for incompatible releases, matrix alignment, and lock consistency. ADK 2.9.2 remains unsupported.
Drive deadlines independently of host scheduling while preserving real
cancellation, durable batch progress and exact missing-chunk recovery.
Keep production deadlines and compression behavior unchanged.

Validate 500ms scheduler pauses, three rejected fault mutations and the
full offline gate on Python 3.10/ADK 1.34 and Python 3.12/ADK 2.2.
Allow Agent(context_compression={"prepare_index": True}) to prepare authorized
sources before query retrieval using separate time and embedding-call budgets.
Reuse completed SQLite batches without consuming the query budget; preserve
source validation, cancellation, no-embedding fallback and Session originals.

Add 14 regression cases to the required context gate and document direct Agent
usage, preparation costs, limits and opt-out behavior.
Give dependency installation a five-minute step limit and retain the existing
three-minute Sidecar test limit within a ten-minute job budget. This prevents
cold installation from cancelling the tests before their execution budget.

Preserve all test commands, coverage requirements and the fail-closed aggregate.
Add a workflow regression to the mandatory release gate.
…nking

Preserve source-backed evidence selection and SQLite recovery. Enable
bounded preparation and ID-only reranking when embedding is configured;
reuse the Agent model for extraction with Ark thinking disabled.

Add a two-turn Agent/Runner tutorial and regression proving request
compression and complete persisted originals. Update developer guides
and include teaching changes in the context gate triggers.

Validation: pre-commit; Python 3.10/ADK 1.34.0 and Python 3.12/ADK 2.2.0
each 1447 passed, 5 skipped; focused suites each 132 passed.
@zyn080302
zyn080302 force-pushed the feat/default-context-compression branch from 9cf4d93 to 7720826 Compare October 9, 2026 09:18

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant