Repository navigation
fix: context window management - #3102
Merged
Merged
Conversation
…ry files ReadFile decoded any file as UTF-8, so a .docx returned ~200KB of zip junk that blew the model context. PDF/DOCX/PPTX/XLSX now go through smssutil.get_document_markdown in the room python process; other binary files return a clear error.
Contributor
✅ Snyk checks have passed. No issues have been found so far.
💻 Catch issues earlier using the plugins for VS Code, JetBrains IDEs, Visual Studio, and Eclipse. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
This PR fixes how agent rooms handle large context, and lets a room recover when one request is too big for the model.
The bug was found in a playground room on Laguna (250k window).
ReadFiledecoded a.docxas UTF-8 and stored a 1.3MB tool result withtokens=0._check_context_budgetindocument_input.py) then rejected every later request.This PR covers four areas:
ReadFileextracts Office/PDF files via Docling and refuses other binary files.Follow-up for the kept-turns size during compaction: #3101.
Changes Made
ReadFile (
PlatformAgentToolHandlers)smssutil.get_document_markdownin the room's Python process.decodeTextOrNullrejects NUL bytes and invalid UTF-8, so binary files return a clear error instead of garbage text.Compaction checks (
SemossAgentHarness,Room,IModelEngine,AbstractModelEngine)estimateContextTokensuses the last measured input tokens plus later messages (chars/4 for those not yet measured).estimate + reserve >= window. The reserve ismax_tokens(taken from the request, else the engine'smaxOutputTokens/MAX_TOKENS, else 16k), plus a 20k margin, capped at 25% of the window.pruneToolsBeforeContinuation). Older tool outputs are pruned from what the model sees; stored history is unchanged.IModelEngine.getMaxTokens()andRoom.markPruneToolsAbove(...).Summary request (
CompactRoomMessagesReactor)max_tokens = clamp(20% of transcript, 2k, min(5% of window, 12k)), plus a word budget in the prompt.<transcript>tags, with the instructions after it. With instructions first and a roughly 130k-token transcript after them, Laguna continued the chat instead of summarizing; the entire summary was one line. Removed the duplicate[SUMMARY]request.Overflow recovery
document_input.py,model_engine_exception.py): newContextBudgetError.ErrorDetailsgainsreason=CONTEXT_OVERFLOWandreason_detail, which records the rule that matched.ContextBudgetError, OpenAI/Azurecontext_length_exceeded, Anthropicrequest_too_large.AskErrorModelEngineResponse(reason,reasonDetail,isContextOverflow()),AskModelEngineResponseandSemossModelEngineException.contextOverflowError(Throwable).HarnessToolExecutor,Room.forkToolResultsWithStubs):SemossAgentHarness.askWithOverflowRecovery):Attachments in pruned history (
MessageUtils)deepCopyForPruningcopied every message above a prune flag throughGSON_FOR_DB, which dropsbase64Data; the transient room folder was lost too. Attachments in those messages reached Python empty ("Attachment ... has no readable data").deepCopykeeps the room links and attachment bytes.How to Test
All of these were run locally against Laguna (
f686bf91-be2c-495e-9e33-1e6c7091d207).initial ask context overflow ... rule=type:ContextBudgetError prunedHistory=truein the log, then the run completes.ReadFileit.rule=text:maximum context length, thenForked tool results ... replaced=1. The model re-reads in parts and the run completes.ReadFileon a.docxreturns extracted markdown. On another binary file it returns an error instead of decoded bytes..mdfile above a pruned message, send a new message.auto compaction completed ... type=SUMMARY tokensBefore=~203k tokensAfter=~82k. The[SUMMARY]is a structured summary, and the answer recalls all codewords and facts.Notes
reason=CONTEXT_OVERFLOWso far.OpenAiJavaEngine,AnthropicJavaEngineandGoogleGenAiJavaEnginedon't yet, so overflow recovery doesn't trigger for them.