Do you really know what your agents are doing when you are not looking?OpenAI hacks HuggingFace

Research

The systems, methods, and experiments that shape how we build and safeguard AI products.

ActionBench: discovering hidden failure modes of agents

Evaluating how agents decide, use tools, and cross boundaries.

Read

PolicyGuard: evaluating runtime control for AI actions

Measuring the real-world performance of action-level guardrails.

Read

Preprint · Security

Coming soon

Tool authorization failure modes

Where permissioning breaks down when models try to act.

Infrastructure

Coming soon

Policy enforcement under pressure

Keeping control decisions fast when agents run at production scale.

Training

Coming soon

Designing human approval loops

Making reinforcement from review practical for production agents.

All systems operational

Works with

  • OpenAI
  • Anthropic
  • Google
  • Microsoft
  • Meta
  • Mistral
  • and self-hosted model deployments
© 2026 Runsphere

Want to deploy AI agents safely?

Control AI behavior in real time with Runsphere