In one paragraph
Two broad problems sit here: checking code an agent writes, and controlling what a running agent may do. automated-security-helper handles the first by orchestrating Bandit, Checkov, Semgrep, Grype and others into one scan pipeline. threat-modeling-mcp-server structures STRIDE threat analysis into nine gated phases. dogwood, dogwood-local-engine, and box address the second: a policy language, a durable incremental evaluation engine, and a sandbox that ties OS isolation to stateful policy enforcement. What counts here: policy languages, sandboxes, and security tooling that limit what an agent may do or check what it builds.
The main approaches
Security review tools
Static analysis, dependency scanning, and IaC review run against code the agent produces or touches. The tool set is pre-assembled; the user points the scanner at a project and reads normalized output. automated-security-helper orchestrates multiple scanners and exposes results over MCP so an agent can query them directly. threat-modeling-mcp-server guides an agent through structured STRIDE analysis of a project, producing a documented threat model with mitigations and an optional check of them against the source.
Policy languages
A formal grammar for writing rules about what an agent may do. Rules are separate from the agent binary and can be validated, compiled, and replayed against event traces. dogwood is a Cedar-derived language that adds temporal predicates so a policy can consult the agent’s recent action history, not just the current request. dogwood-local-engine evaluates those policies incrementally, checkpointing state so verdicts survive process restarts.
Sandboxes
OS-level isolation wraps the agent process. File and network access is denied by default; explicit grants in configuration open specific paths or destinations. box combines macOS sandbox primitives with Dogwood policy evaluation and an egress proxy that injects credentials into permitted outbound calls without exposing secrets to the agent.
Map of the theme
flowchart LR t["Agent security and policy"] t --> f1["Policy languages"] t --> f2["Sandboxes"] t --> f3["Security review tools"] f1 --> r1["dogwood-policy/dogwood"] f1 --> r2["dogwood-policy/dogwood-local-engine"] f2 --> r3["strands-agents/box"] f3 --> r4["awslabs/automated-security-helper"] f3 --> r5["awslabs/threat-modeling-mcp-server"]
Where the new ideas are
Temporal policy reasoning. dogwood adds formerly, since, and windowed aggregation predicates to Cedar’s permit/forbid model. A policy can block a network call because of a file read that happened earlier in the same session, which standard Cedar rules evaluated against the current request alone cannot express.
Durable incremental evaluation. dogwood-local-engine processes events one at a time and checkpoints derived monitor state to disk. The engine resumes correctly after a crash without replaying the full event history.
Credential injection outside the agent. box injects secrets into permitted egress calls at the proxy layer. The agent process never holds the credentials.
MCP as a security interface. automated-security-helper exposes its scan pipeline as an MCP server so an AI coding agent can trigger scans, diff results, and draft suppressions without leaving its tool-call loop.
Side by side
| Repo | Approach | Output formats | MCP support | Temporal/stateful policy | OS-level isolation |
|---|---|---|---|---|---|
| awslabs/automated-security-helper | Security review tools | SARIF, JUnit XML, HTML, Markdown, CSV, JSON | Yes | Not stated | Not stated |
| awslabs/threat-modeling-mcp-server | Security review tools | Markdown report, Threat Composer .tc.json | Yes | Not stated | Not stated |
| dogwood-policy/dogwood | Policy languages | Lowered Cedar policies, per-event verdict text | Not stated | Yes, via formerly/since/windowed predicates | Not stated |
| dogwood-policy/dogwood-local-engine | Policy languages | In-process verdicts, demo Unix-socket daemon | Not stated | Yes, incremental with disk checkpoints | Not stated |
| strands-agents/box | Sandboxes | OTLP JSON telemetry | Yes (MCP broker) | Yes, via embedded Dogwood engine | Yes (macOS Apple silicon) |
How the idea moved
flowchart LR n1["started May 2022<br/>awslabs/automated-security-helper"] n2["started Dec 2025<br/>awslabs/threat-modeling-mcp-server"] n3["started Jul 2026<br/>dogwood-policy/dogwood"] n4["started Aug 2026<br/>dogwood-policy/dogwood-local-engine"] n5["started Oct 2026<br/>strands-agents/box"] n1 --> n2 --> n3 --> n4 --> n5
- started May 2022 · automated-security-helper · Adds: Orchestrates Bandit, Checkov, Semgrep, Grype and others into one scan pipeline with normalized SARIF/JSON/HTML output; its v3 rewrite adds an MCP server so AI coding agents can trigger scans and read results directly.
- started Dec 2025 · threat-modeling-mcp-server · Adds: Walks an AI agent through nine-phase STRIDE threat modeling with completion checks and exports a Threat Composer JSON file plus a Markdown report.
- started Jul 2026 · dogwood · Adds: Defines a Cedar-derived policy language with temporal predicates (formerly, since, windowed aggregations) that let rules consult an agent’s past action history, and ships a Rust reference parser, interpreter, and CLI.
- started Aug 2026 · dogwood-local-engine · Adds: Evaluates Dogwood policies incrementally against a stream of agent events, checkpointing derived monitor state to disk so verdicts survive process restarts.
- started Oct 2026 · box · Adds: Combines macOS OS-level sandboxing with an embedded Dogwood policy engine and an egress proxy that injects credentials into permitted outbound calls, keeping secrets outside the agent process.
Easily confused
This is not general agent orchestration. These repos do not route tasks between agents or manage agent memory. They constrain what an agent may do (policy, sandbox) or audit what it has built (scanner, threat model).
This is not authentication or identity management. The policy engines here assume an agent identity is already established; they decide what that identity may do next, based on rules and history.
The Dogwood repos are not a full agent framework. They are a policy language and its evaluation engine. A sandbox, such as box, embeds them to enforce rules, but the language and engine repos have no agent execution logic of their own.
Gaps nobody has filled
- Nothing here describes one policy set and event history shared across several separate agent sandboxes running at once. Box shares one history across the interpreters and gateways of a single box.
- Nothing here measures the coverage or completeness of a policy set against a known threat taxonomy.
- The OS-level sandbox in strands-agents/box targets macOS on Apple silicon only; Linux support is listed as planned, not present.
- Nothing here provides a standard format for exchanging policy artifacts between the scanner and the policy language layers.