Guardrails
Every MCP tool call passes through a guardrail evaluation layer before execution. This layer, called the Policy Enforcement Point (PEP) and Policy Decision Point (PDP), evaluates your configured rules deterministically against the tool call’s arguments and returns a decision: allow, deny, require approval, or mask fields from the result.
Guardrails are a deterministic safety mechanism between the agent and your MCP servers. They are not probabilistic and not model-based. Every rule is a deterministic condition that either matches or does not match.
The primary safety layer sits below the guardrails you author: an assurance gate derived from each tool’s risk classification. It runs on every call, before any of your guardrail rules, and it is described first below. The guardrail rules you configure (deny, approval gate, mask, escalate) layer on top of it.
How evaluation works
Section titled “How evaluation works”When the agent calls a tool, MCP Hub evaluates the call in this order:
- Account isolation. Hard deny if the connection does not belong to the requesting account.
- Rate limit check. Deny if the account has exceeded the rate limit for this tool.
- Assurance gate. Resolve the tool’s required assurance level from its risk classification. An unclassified tool is hard-denied. A call below the required assurance level is denied. See “Risk and the assurance gate” below.
- Guardrail rules. Evaluate all enabled guardrails that match the tool name. Deny rules take precedence over all other types (forbid-overrides-permit).
- Default risk policy. If no guardrail matched and the tool is classified as
destructive, require approval automatically. - Default allow. If nothing blocked the call, allow it.
On any evaluation error (malformed condition, missing field, unexpected type), the system denies the call. This is fail-closed by design.
Policy types
Section titled “Policy types”Each guardrail has a policyType that determines what happens when the rule’s condition matches.
| Policy type | Execution | Behavior |
|---|---|---|
deny | Blocked | Tool call is rejected immediately. The agent receives an error and cannot retry. |
approval_gate | Blocked | Tool call is paused. An approval request is created and the conversation workflow waits for a human to approve or reject. |
mask | Allowed | Tool call executes normally. Specified fields are redacted from the result before it reaches the agent. |
escalate | Allowed | Tool call executes normally. After execution, the result is evaluated against the condition. If it matches, an escalation signal is attached for post-execution review. |
When multiple guardrails match the same tool call, the evaluation follows forbid-overrides-permit: a single deny match blocks the call regardless of any allow or approval_gate rules. An approval_gate match blocks the call even if other rules would allow it.
Managing guardrails
Section titled “Managing guardrails”All guardrail endpoints require dashboard authentication via Authorization: Bearer <token>. Guardrails are scoped to an MCP connection: each guardrail applies to a specific tool (or all tools) on a specific connection.
List guardrails
Section titled “List guardrails”Returns all guardrails for a connection.
curl -X GET https://api.ohallo.eu/api/mcp-connections/{connectionId}/guardrails \ -H "Authorization: Bearer <token>"[ { "id": "a1b2c3d4-e5f6-7890-abcd-ef1234567890", "tenantId": "t1234567-abcd-efgh-ijkl-mnopqrstuvwx", "mcpConnectionId": "c9876543-dcba-fedc-ba98-765432109876", "toolName": "submit_return", "policyType": "approval_gate", "config": { "condition": "args.refund_amount > 500", "description": "Require approval for returns over 500 EUR" }, "enabled": true, "createdAt": "2026-04-10T14:30:00.000Z", "updatedAt": "2026-04-10T14:30:00.000Z" }]Create a guardrail
Section titled “Create a guardrail”Creates a new guardrail on a connection. If a condition is provided, it is validated as a filtrex expression at creation time. Invalid expressions return a 422 error.
curl -X POST https://api.ohallo.eu/api/mcp-connections/{connectionId}/guardrails \ -H "Authorization: Bearer <token>" \ -H "Content-Type: application/json" \ -d '{ "toolName": "submit_return", "policyType": "approval_gate", "condition": "args.refund_amount > 500", "config": { "description": "Require approval for returns over 500 EUR" }, "enabled": true }'Request body:
| Field | Type | Required | Description |
|---|---|---|---|
toolName | string | Yes | The tool this guardrail applies to. Use * to match all tools on the connection. |
policyType | string | Yes | One of: approval_gate, deny, mask, escalate. |
condition | string | No | A filtrex expression evaluated against the tool call arguments. If omitted, the guardrail matches every call to the tool. |
config | object | No | Additional configuration. Use description to explain the rule, maskFields for mask type, escalationType for escalate type. |
enabled | boolean | No | Defaults to true. Set to false to disable the guardrail without deleting it. |
Response (201):
{ "id": "a1b2c3d4-e5f6-7890-abcd-ef1234567890", "tenantId": "t1234567-abcd-efgh-ijkl-mnopqrstuvwx", "mcpConnectionId": "c9876543-dcba-fedc-ba98-765432109876", "toolName": "submit_return", "policyType": "approval_gate", "config": { "condition": "args.refund_amount > 500", "description": "Require approval for returns over 500 EUR" }, "enabled": true, "createdAt": "2026-04-10T14:30:00.000Z", "updatedAt": "2026-04-10T14:30:00.000Z"}Error (422), invalid condition:
{ "error": "Invalid condition expression", "details": "Unexpected token at position 12"}Update a guardrail
Section titled “Update a guardrail”Updates an existing guardrail. Only the fields you include in the request body are changed.
curl -X PATCH https://api.ohallo.eu/api/mcp-call-guardrails/{guardrailId} \ -H "Authorization: Bearer <token>" \ -H "Content-Type: application/json" \ -d '{ "condition": "args.refund_amount > 1000", "config": { "description": "Require approval for returns over 1000 EUR" } }'Request body (all fields optional):
| Field | Type | Description |
|---|---|---|
toolName | string | Change the tool this guardrail applies to. |
policyType | string | Change the policy type. |
condition | string | Change the condition expression. Validated at update time. |
config | object | Replace the configuration object. |
enabled | boolean | Enable or disable the guardrail. |
Response (200): The updated guardrail object.
Error (404): Guardrail not found or does not belong to your account.
Delete a guardrail
Section titled “Delete a guardrail”curl -X DELETE https://api.ohallo.eu/api/mcp-call-guardrails/{guardrailId} \ -H "Authorization: Bearer <token>"Response (200):
{ "ok": true}Bulk upsert guardrails
Section titled “Bulk upsert guardrails”Replaces all guardrails for a connection at once. Existing guardrails on the connection are deleted and the provided list is inserted atomically. This is the endpoint the oHallo dashboard uses when saving guardrail configuration.
curl -X PUT https://api.ohallo.eu/api/mcp-connections/{connectionId}/policies \ -H "Authorization: Bearer <token>" \ -H "Content-Type: application/json" \ -d '{ "guardrails": [ { "toolName": "submit_return", "policyType": "approval_gate", "config": { "condition": "args.refund_amount > 500", "description": "Require approval for returns over 500 EUR" }, "enabled": true }, { "toolName": "delete_customer", "policyType": "deny", "config": { "description": "Never allow customer deletion via the agent" } } ] }'Response (200): The full list of created guardrail objects.
Conditions
Section titled “Conditions”Guardrail conditions use filtrex, a sandboxed expression language that evaluates against the tool call arguments. No eval or function constructors are used. Expressions are compiled to a safe AST.
The oHallo dashboard provides a visual condition builder that compiles to filtrex automatically. Operators enter a field name, select an operator, and type a value, and the args. prefix and expression syntax are handled by the UI. The API accepts the compiled filtrex expression directly.
Syntax
Section titled “Syntax”Conditions access tool call arguments through the args namespace using dot notation:
| Expression | Description |
|---|---|
args.refund_amount > 500 | Matches when the refund amount exceeds 500 |
args.quantity > 100 | Matches when quantity exceeds 100 |
args.country == "US" | Matches when the country is US |
args.priority == "critical" | Matches when priority is critical |
You can combine conditions with logical operators:
| Operator | Example |
|---|---|
and | args.refund_amount > 500 and args.currency == "EUR" |
or | args.priority == "critical" or args.priority == "urgent" |
not | not (args.is_verified == 1) |
Comparisons: ==, !=, <, >, <=, >=.
Arithmetic: +, -, *, /.
No condition
Section titled “No condition”If a guardrail has no condition (the field is omitted), it matches every call to the specified tool. Use this for blanket rules like “never allow deletion” or “always require approval for this tool.”
Evaluation errors
Section titled “Evaluation errors”If a condition references a field that does not exist in the tool call arguments, or if the expression is malformed, the evaluation fails and the tool call is denied. This is fail-closed: a broken condition blocks execution rather than silently allowing it.
Risk and the assurance gate
Section titled “Risk and the assurance gate”The assurance gate is the primary safety layer. It runs on every tool call, before any guardrail rule you author, and it is derived deterministically from the tool’s risk classification. Nothing you configure can weaken it.
Five risk classifications
Section titled “Five risk classifications”Every tool on an MCP connection has a risk classification. There are five:
| Classification | Meaning |
|---|---|
read | Public or non-personal data. |
read_pii | Returns personal data. |
write | State-changing action. |
write_pii | State-changing action that also touches personal data. Handled the same as write; the distinct value exists for audit clarity. |
destructive | Irreversible or financially material. |
The classification maps to a required assurance level
Section titled “The classification maps to a required assurance level”Each risk classification carries a required assurance level (the “floor”). A call whose resolved assurance is below the tool’s floor is denied with assurance_too_low. The floors are fixed by the platform:
| Classification | Required assurance floor |
|---|---|
read | anonymous (no identification required) |
read_pii | identified |
write | identified |
write_pii | identified |
destructive | high-assurance |
You can raise the required assurance above the floor for a specific tool with a per-tool override, but you can never lower it below the floor. This mapping is enforced identically by MCP Hub, the API, and the dashboard: what is displayed can never diverge from what is enforced.
Unclassified tools are denied
Section titled “Unclassified tools are denied”If a tool has no risk classification, it is hard-denied on every call and is non-callable until you classify it. An unknown capability is treated as maximum risk. There is no “no classification, so allow” fallback and no default to write. Classify every tool before assigning it to an agent.
Setting risk classifications
Section titled “Setting risk classifications”curl -X PATCH https://api.ohallo.eu/api/mcp-connections/{connectionId}/tool-risk-classifications \ -H "Authorization: Bearer <token>" \ -H "Content-Type: application/json" \ -d '{ "classifications": { "get_order": "read", "get_customer": "read_pii", "update_order": "write", "delete_order": "destructive", "submit_return": "write", "cancel_subscription": "destructive" } }'Response (200): The full connection object with the updated classifications.
Risk classifications and guardrails work together. The assurance gate runs first: an unclassified tool is denied, and a call below the tool’s assurance floor is denied, before any guardrail is consulted. A read tool with a deny guardrail is still blocked when the condition matches. A destructive tool with an existing approval_gate guardrail uses that guardrail’s condition; the default “require approval” only applies when no guardrail matched at all.
Approval lifecycle
Section titled “Approval lifecycle”When a tool call triggers an approval_gate guardrail (or the default destructive-tool policy), the following sequence occurs:
1. Tool call returns approval_required
Section titled “1. Tool call returns approval_required”The agent receives a response indicating that the tool call needs human approval:
{ "success": false, "error_type": "approval_required", "reason": "Require approval for returns over 500 EUR", "policy_id": "a1b2c3d4-e5f6-7890-abcd-ef1234567890", "approval_request_id": "req-5678-abcd-efgh-ijkl-9012", "expires_at": "2026-04-11T14:30:00.000Z", "policy_decision": "require_approval"}2. Workflow pauses
Section titled “2. Workflow pauses”The conversation workflow creates an attention item visible in the oHallo dashboard and waits for a human decision. The customer is not sent a reply while approval is pending.
3. Human reviews and decides
Section titled “3. Human reviews and decides”A team member reviews the pending tool call in the dashboard. They see the tool name, arguments, the guardrail that triggered, and the reason. They approve or reject via the dashboard UI or the API:
curl -X POST https://api.ohallo.eu/api/conversations/{conversationId}/approve \ -H "Authorization: Bearer <token>" \ -H "Content-Type: application/json" \ -d '{ "approvalRequestId": "req-5678-abcd-efgh-ijkl-9012", "approved": true, "reviewerNote": "Confirmed with warehouse, return is valid" }'Request body:
| Field | Type | Required | Description |
|---|---|---|---|
approvalRequestId | string | Yes | The ID from the approval_required response. |
approved | boolean | Yes | true to approve, false to reject. |
reviewerNote | string | No | Optional note explaining the decision. Stored for audit. |
Response (200):
{ "ok": true}4. Workflow resumes
Section titled “4. Workflow resumes”After approval, the conversation workflow resumes. The agent retries the tool call with the same arguments. MCP Hub finds the approved request record and allows the call through without triggering the guardrail again.
After rejection, the workflow resumes and the agent is informed that the action was denied. It composes a response to the customer accordingly.
Approval expiry
Section titled “Approval expiry”Approval requests expire after 24 hours. If no decision is made within that window, the request expires and the workflow treats it as a rejection.
Idempotent retries
Section titled “Idempotent retries”If the agent retries a tool call that already has a pending approval request with the same arguments, MCP Hub returns the existing request instead of creating a duplicate. Arguments are matched by a SHA-256 hash of the sorted argument keys and values.
Examples
Section titled “Examples”Block a tool entirely
Section titled “Block a tool entirely”Prevent the agent from ever calling delete_customer, regardless of arguments:
curl -X POST https://api.ohallo.eu/api/mcp-connections/{connectionId}/guardrails \ -H "Authorization: Bearer <token>" \ -H "Content-Type: application/json" \ -d '{ "toolName": "delete_customer", "policyType": "deny", "config": { "description": "Customer deletion is never allowed via the agent" } }'Require approval above a threshold
Section titled “Require approval above a threshold”Require human approval for refunds over 500 EUR:
curl -X POST https://api.ohallo.eu/api/mcp-connections/{connectionId}/guardrails \ -H "Authorization: Bearer <token>" \ -H "Content-Type: application/json" \ -d '{ "toolName": "submit_return", "policyType": "approval_gate", "condition": "args.refund_amount > 500", "config": { "description": "Require approval for returns over 500 EUR" } }'Mask sensitive fields in results
Section titled “Mask sensitive fields in results”Allow the tool call but redact Social Security numbers and bank account details from the result:
curl -X POST https://api.ohallo.eu/api/mcp-connections/{connectionId}/guardrails \ -H "Authorization: Bearer <token>" \ -H "Content-Type: application/json" \ -d '{ "toolName": "get_customer", "policyType": "mask", "config": { "description": "Redact sensitive financial data", "maskFields": ["ssn", "bank_account", "credit_card"] } }'Flag high-value operations for review
Section titled “Flag high-value operations for review”Allow the tool call but flag orders above 10,000 EUR for post-execution review:
curl -X POST https://api.ohallo.eu/api/mcp-connections/{connectionId}/guardrails \ -H "Authorization: Bearer <token>" \ -H "Content-Type: application/json" \ -d '{ "toolName": "create_order", "policyType": "escalate", "condition": "args.total > 10000", "config": { "description": "Flag high-value orders for review", "escalationType": "high_value_order" } }'Next steps
Section titled “Next steps”- Build an MCP Server: create the server your guardrails will protect
- Tool Schema: define the input schemas that guardrail conditions evaluate against