Guardrails ✦ 85outlineNot a production recipe — calibrate thresholds

Jailbreak / injection input screen

Who: LLM app owners. Steps: State=user message → Noul jailbreak/injection + Score harm → pass/review/block. Expected effect: Fail closed on high hazard.

1

Treat this as an outline — adapt state and questions to your data.

2

Implement in code — call System One / Jev; compose answers yourself.

3

Gate on confidence — act, confirm, or escalate before side effects.

Outline sketch

Prompt
# Jailbreak / injection input screen
State: relevant software state for this pattern.
Ask: Choice / Score / Noul questions that close the decision space.
Then: compose the route or action in code; escalate when confidence is low.

Needs access to: typesafe-sdk

Who it's for

LLM app owners

Steps / how it's set up

Steps

  1. State=user message.
  2. Noul jailbreak/injection + Score harm.
  3. pass/review/block.

Sources

Metrics and demo claims are author-reported or cookbook-reported unless you measure them yourself. Calibrate on your data.

Expected effect

Fail closed on high hazard

Unofficial outline for learning. Paraphrased from public docs and tutorials — not a production recipe. Review sources before you automate anything.