Skip to content

feat:jev guardrail - #3096

Draft
kunal0137 wants to merge 10 commits into
SEMOSS:devfrom
kunal0137:feat/jev-guardrail
Draft

kunal0137 wants to merge 10 commits into
SEMOSS:devfrom
kunal0137:feat/jev-guardrail

Conversation

@kunal0137

Copy link
Copy Markdown
Collaborator

Description

Changes Made

How to Test

  1. Steps to reproduce/test the behavior
  2. Expected outcomes

Notes

One engine class evaluates text against a configurable policy criterion
by asking a configured Jev (TypeSafe) model engine typed questions and
mapping the structured answer to a pass/block verdict. Noul (binary
criterion) and choice question types are supported through the same
implementation; different policies change configuration only.

- SMSS: MODEL_ENGINE_ID, QUESTION_INSTRUCTIONS, optional QUESTION_TYPE,
  QUESTION_CRITERIA/VIOLATION_CHOICES (choice), VIOLATION_DIRECTION
  (noul), CONFIDENCE_THRESHOLD (inclusive boundary), LOW_CONFIDENCE_
  VERDICT (default BLOCK), FAIL_OPEN (default false), BLOCKED_MESSAGE,
  TIMEOUT_SECONDS; per-call overrides via pipeline directParameters
- Structured verdict details distinguish PASS / VIOLATION /
  INDETERMINATE / ERROR outcomes with criterion, typed answer,
  confidence vs threshold, and reason
- Judge resolution enforces model access (userCanViewEngine) and the
  judge engine's usage controls run inside evaluate(); recursion into
  the same guardrail from the judge's own pipeline is detected and
  blocked per thread, and a guardrail mounted on its judge's pipeline
  fails immediately
- Evaluation errors (missing/unreachable/malformed judge, missing user
  context, invalid overrides) follow the configured failure policy and
  block by default - they never become silent passes
- Registered as GuardrailTypeEnum.EMBEDDED_JEV_POLICY; docs include two
  pipeline examples showing two policies from one engine
A passthrough route (Ollama/Anthropic chat endpoints) hands the guarded
engine an InputMessage whose text part is the pixel placeholder while the
real conversation rides in the full_prompt parameter, and the engine
layer replaces its input from full_prompt just before inference. A
guardrail screening arg0 was therefore judging the placeholder.

PromptGuardrailEngine.fullPromptText(InputMessage) reads the full_prompt
parameter (JSON or list of chat-shaped entries); the last user turn's
contents are what an input guardrail screens, falling back to all
contents when no user turn exists. GuardrailValueReader.messageText
prefers it over the placeholder text part, so all guardrail engines
benefit. chat-shaped message maps inside a plain parameter map resolve
through extractText as well.

17 Jev + 6 PolicyCompliance unit tests pass.
deepset/prompt-injections, 546 rows through the real Jev API: the shipped
criterion (v3) catches 158/203 injections with zero false positives at
the default 0.5 threshold. Iterations recorded in README.
On the streaming routes the model's partial output drains to the client
while the guarded method is still running, so an output guardrail reviews
the completed response after chunks are already emitted; a violation
cannot be recalled. Say so explicitly instead of implying delivered
content is protected.
@kunal0137 kunal0137 changed the title Feat/jev guardrail feat:jev guardrail Oct 7, 2026
@snyk-io

snyk-io Bot commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

✅ Snyk checks have passed. No issues have been found so far.

Status Scan Engine Critical High Medium Low Total (0)
✅ Open Source Security 0 0 0 0 0 issues

💻 Catch issues earlier using the plugins for VS Code, JetBrains IDEs, Visual Studio, and Eclipse.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant