Repository navigation
feat:jev guardrail - #3096
Draft
kunal0137 wants to merge 10 commits into
Draft
feat:jev guardrail#3096kunal0137 wants to merge 10 commits into
kunal0137 wants to merge 10 commits into
Conversation
One engine class evaluates text against a configurable policy criterion by asking a configured Jev (TypeSafe) model engine typed questions and mapping the structured answer to a pass/block verdict. Noul (binary criterion) and choice question types are supported through the same implementation; different policies change configuration only. - SMSS: MODEL_ENGINE_ID, QUESTION_INSTRUCTIONS, optional QUESTION_TYPE, QUESTION_CRITERIA/VIOLATION_CHOICES (choice), VIOLATION_DIRECTION (noul), CONFIDENCE_THRESHOLD (inclusive boundary), LOW_CONFIDENCE_ VERDICT (default BLOCK), FAIL_OPEN (default false), BLOCKED_MESSAGE, TIMEOUT_SECONDS; per-call overrides via pipeline directParameters - Structured verdict details distinguish PASS / VIOLATION / INDETERMINATE / ERROR outcomes with criterion, typed answer, confidence vs threshold, and reason - Judge resolution enforces model access (userCanViewEngine) and the judge engine's usage controls run inside evaluate(); recursion into the same guardrail from the judge's own pipeline is detected and blocked per thread, and a guardrail mounted on its judge's pipeline fails immediately - Evaluation errors (missing/unreachable/malformed judge, missing user context, invalid overrides) follow the configured failure policy and block by default - they never become silent passes - Registered as GuardrailTypeEnum.EMBEDDED_JEV_POLICY; docs include two pipeline examples showing two policies from one engine
A passthrough route (Ollama/Anthropic chat endpoints) hands the guarded engine an InputMessage whose text part is the pixel placeholder while the real conversation rides in the full_prompt parameter, and the engine layer replaces its input from full_prompt just before inference. A guardrail screening arg0 was therefore judging the placeholder. PromptGuardrailEngine.fullPromptText(InputMessage) reads the full_prompt parameter (JSON or list of chat-shaped entries); the last user turn's contents are what an input guardrail screens, falling back to all contents when no user turn exists. GuardrailValueReader.messageText prefers it over the placeholder text part, so all guardrail engines benefit. chat-shaped message maps inside a plain parameter map resolve through extractText as well. 17 Jev + 6 PolicyCompliance unit tests pass.
deepset/prompt-injections, 546 rows through the real Jev API: the shipped criterion (v3) catches 158/203 injections with zero false positives at the default 0.5 threshold. Iterations recorded in README.
This reverts commit 12f9b5f.
)" This reverts commit 34623f2.
On the streaming routes the model's partial output drains to the client while the guarded method is still running, so an output guardrail reviews the completed response after chunks are already emitted; a violation cannot be recalled. Say so explicitly instead of implying delivered content is protected.
Contributor
✅ Snyk checks have passed. No issues have been found so far.
💻 Catch issues earlier using the plugins for VS Code, JetBrains IDEs, Visual Studio, and Eclipse. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Changes Made
How to Test
Notes