Skip to content

fix: context window management - #3102

Merged
kunal0137 merged 4 commits into
devfrom
fix-context-window-management
Oct 7, 2026
Merged

kunal0137 merged 4 commits into
devfrom
fix-context-window-management

Conversation

@kunal0137

@kunal0137 kunal0137 commented Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator

Description

This PR fixes how agent rooms handle large context, and lets a room recover when one request is too big for the model.

The bug was found in a playground room on Laguna (250k window).

  • Trigger: ReadFile decoded a .docx as UTF-8 and stored a 1.3MB tool result with tokens=0.
  • Effect: the Python preflight (_check_context_budget in document_input.py) then rejected every later request.
  • Why nothing caught it: auto compaction only ran at the start of a run, never between tool calls. And when a request was too big, nothing recovered the room.

This PR covers four areas:

  1. Prevention: ReadFile extracts Office/PDF files via Docling and refuses other binary files.
  2. Compaction: the context check now runs every turn, with a better size estimate and a fixed output reserve. The summary prompt was also fixed.
  3. Overflow recovery: when a request is too large for the model, the oldest or largest tool output is replaced and the request is retried, instead of leaving the room stuck.
  4. Attachment bug: pruned history no longer drops attachment data.

Follow-up for the kept-turns size during compaction: #3101.

Changes Made

ReadFile (PlatformAgentToolHandlers)

  • pdf/docx/pptx/xlsx are converted to markdown through smssutil.get_document_markdown in the room's Python process.
  • decodeTextOrNull rejects NUL bytes and invalid UTF-8, so binary files return a clear error instead of garbage text.

Compaction checks (SemossAgentHarness, Room, IModelEngine, AbstractModelEngine)

  • Estimate: estimateContextTokens uses the last measured input tokens plus later messages (chars/4 for those not yet measured).
  • Trigger: estimate + reserve >= window. The reserve is max_tokens (taken from the request, else the engine's maxOutputTokens/MAX_TOKENS, else 16k), plus a 20k margin, capped at 25% of the window.
  • Mid-run: the check also runs after each tool batch (pruneToolsBeforeContinuation). Older tool outputs are pruned from what the model sees; stored history is unchanged.
  • Run start: if the latest message is unanswered tool results left by a failed run, prune under it so the room can continue.
  • New API: IModelEngine.getMaxTokens() and Room.markPruneToolsAbove(...).

Summary request (CompactRoomMessagesReactor)

  • Size cap: summary max_tokens = clamp(20% of transcript, 2k, min(5% of window, 12k)), plus a word budget in the prompt.
  • Prompt order: the transcript now comes first inside <transcript> tags, with the instructions after it. With instructions first and a roughly 130k-token transcript after them, Laguna continued the chat instead of summarizing; the entire summary was one line. Removed the duplicate [SUMMARY] request.

Overflow recovery

  • Python (document_input.py, model_engine_exception.py): new ContextBudgetError. ErrorDetails gains reason=CONTEXT_OVERFLOW and reason_detail, which records the rule that matched.
    • Structured signals are checked first: ContextBudgetError, OpenAI/Azure context_length_exceeded, Anthropic request_too_large.
    • Then wording, for vLLM, Anthropic token limits, Gemini and Bedrock.
  • Java plumbing: AskErrorModelEngineResponse (reason, reasonDetail, isContextOverflow()), AskModelEngineResponse and SemossModelEngineException.contextOverflowError(Throwable).
  • During a run (HarnessToolExecutor, Room.forkToolResultsWithStubs):
    • The tool results are copied to a new branch (the original is kept), with the largest results replaced by a note until the request fits.
    • The note asks the model to re-run narrower or read in parts. Then the call is retried once.
  • First call of a run (SemossAgentHarness.askWithOverflowRecovery):
    • Prune tool output from the history and retry once.
    • If there's nothing to prune, return a clear "message too large" error.
    • The input is only stored on success, so the room never gets stuck.

Attachments in pruned history (MessageUtils)

  • The bug: deepCopyForPruning copied every message above a prune flag through GSON_FOR_DB, which drops base64Data; the transient room folder was lost too. Attachments in those messages reached Python empty ("Attachment ... has no readable data").
  • The fix: only messages with tool parts are copied. deepCopy keeps the room links and attachment bytes.

How to Test

All of these were run locally against Laguna (f686bf91-be2c-495e-9e33-1e6c7091d207).

  1. A room stuck on a huge tool result. Open the room and send any message.
    • Expected: initial ask context overflow ... rule=type:ContextBudgetError prunedHistory=true in the log, then the run completes.
    • Observed: the follow-up runs were 28-35k input tokens and the run COMPLETED.
  2. A tool result too big mid-run. Put a large text file (about 3MB) in a room and ask the agent to ReadFile it.
    • Expected: the provider rejects the request (vLLM: "maximum context length is 262144"). The log shows rule=text:maximum context length, then Forked tool results ... replaced=1. The model re-reads in parts and the run completes.
    • Check: a follow-up in the same room still works (about 18k input tokens).
  3. Binary file. ReadFile on a .docx returns extracted markdown. On another binary file it returns an error instead of decoded bytes.
  4. Attachments plus pruning. In a room with an attached .md file above a pruned message, send a new message.
    • Expected: no "has no readable data" error.
  5. Compaction. Paste 5 parts of about 43k tokens each into a new room, each with a codeword, and state a couple of facts in part 1. Then ask a recall question.
    • Expected: auto compaction completed ... type=SUMMARY tokensBefore=~203k tokensAfter=~82k. The [SUMMARY] is a structured summary, and the answer recalls all codewords and facts.

Notes

  • Engine coverage: only the Python-backed engines set reason=CONTEXT_OVERFLOW so far. OpenAiJavaEngine, AnthropicJavaEngine and GoogleGenAiJavaEngine don't yet, so overflow recovery doesn't trigger for them.
  • Resume path: resuming a run (after an approval) has no overflow retry yet.
  • Kept turns: compaction still keeps the last 2 turns word for word, whatever their size (about 87k tokens in the test above). The move to a fixed token budget, and how other harnesses handle this, is tracked in Agent compaction: cap the kept recent turns with a token budget instead of the last 2 turns #3101.
  • Estimate accuracy: chars/4 overestimates on repetitive text (about 290k estimated vs 211k real in one test), so pruning can start a little early. Normal prose was close (about 4.4 chars/token on Laguna).
  • Tests: no JUnit tests added, per repo testing guidance. These changes were verified against a local Monolith.

…ry files

ReadFile decoded any file as UTF-8, so a .docx returned ~200KB of zip junk
that blew the model context. PDF/DOCX/PPTX/XLSX now go through
smssutil.get_document_markdown in the room python process; other binary
files return a clear error.
@snyk-io

snyk-io Bot commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

✅ Snyk checks have passed. No issues have been found so far.

Status Scan Engine Critical High Medium Low Total (0)
✅ Open Source Security 0 0 0 0 0 issues

💻 Catch issues earlier using the plugins for VS Code, JetBrains IDEs, Visual Studio, and Eclipse.

@kunal0137
kunal0137 marked this pull request as ready for review October 7, 2026 22:57
@kunal0137
kunal0137 requested a review from a team as a code owner October 7, 2026 22:57
@kunal0137
kunal0137 merged commit ba22b60 into dev Oct 7, 2026
5 checks passed
@kunal0137
kunal0137 deleted the fix-context-window-management branch October 7, 2026 23:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant