AI evaluation · governance · assurance

Understand the risk.
Know what to
do next.

We test AI systems to find where they fail, then work with your team to decide what needs to change—and how to know it worked.

Founded by Nathan Heath · Decision scientist & AI safety researcher ↗

SYNTONY / IN PRACTICE
Fine blue waveforms converge into a shared rhythm, reflecting evaluation, governance, and assurance working together.
Separate disciplines.
A shared understanding.
FIG. 01
RESONANCE
01 EvaluateFind the failure. Keep the evidence.
02 GovernDecide who acts and what changes.
03 AssureTest again as the system evolves.
TWO INTERACTIVE CHALLENGES · CLOSE THE LOOP

Find the fracture.
Follow the consequences.

Two agents. Three escalating rounds. Follow a refund cascade or an evidence echo from the first failure to its downstream consequences.

Challenge Atlas in customer operations or Prism in research. Compare hosted models, choose a control, track delayed effects, and export evidence to GitHub.

Enter the challenge Synthetic worlds · Guided practice + hosted inference
THE REASON WE EXIST

A finding is only useful
if someone can act on it.

A red-team report can show you what went wrong. The harder work comes next: understanding why it matters, deciding who needs to act, and checking the response. We help your technical and governance teams work through that together.

01 / The work, made concrete

What happens after
we find a problem?

Follow a finding from the first test to the next decision. These illustrative examples show the record we help your team build.

Finding

A model follows hostile instructions embedded in retrieved content.

Evidence

Trace, prompt, tool call, affected privilege, and replayable test case.

Owner

System security owns remediation. The program owner owns release conditions.

Control

Content isolation, tool-call approval, and a retrieval allowlist.

Decision

Hold privileged tool use from untrusted retrieved content until the control passes.

Re-test

Replay the attack after remediation. Run it again whenever the model or retrieval layer changes materially.

02 / How we can help

Start with the question
you need to answer.

You may be preparing a release, responding to a finding, or reviewing a system already in use. We’ll agree on the question, the evidence needed, and a useful scope of work.

01 / EVALUATE

Where does it fail?

We test models, agents, and workflows under realistic and adversarial conditions. You get documented failures, the evidence behind them, and test cases your team can run again.

Red teamingEvaluation
Explore evaluation →
02 / GOVERN

What should change?

We help the people responsible for the system turn findings into controls, clear responsibilities, and release decisions. The process needs to work for the team using it.

ControlsDecision support
Explore governance →
03 / ASSURE

Does the fix still hold?

Models, data, and tools change. We keep the tests and decision record current, so your next review starts with evidence about the system you actually have.

Continuous evaluationRe-testing
Explore assurance →
03 / From the research bench

Tools we’re building.

Three projects exploring how to keep AI findings, decisions, and follow-up work connected. Each is still in development.

In development

Governance Lag Monitor

Tracks whether public frontier-AI findings lead to documented controls or decisions, and how long that response takes.

Explore the Monitor →
In development · Public preview

Resonance

Maps attack paths through a customer’s AI system. Each path stays linked to the test record, the responsible team, the response, and the next test.

Open the risk map →
In development · Design-partner pilots

GovTune AI

Drafts governance records from red-team traces, enabling organizations to assign responsibility and plan mitigations while preserving the decision workflows.

Explore GovTune AI →
Who we work with

Different institutions.
Real decisions about AI.

We work with AI labs, companies, public agencies, universities, and civil-society organizations. The questions and constraints differ. We shape the work around yours.

04 / The person behind Syntony

Research is technical.
The work is personal.

Nathan Heath Founder, Syntony

Nathan founded Syntony because good evaluation work too often stalls between the team that finds a problem and the people who can change the system. His work brings those people and their decisions into the same conversation.

Nathan red-teams frontier models for OpenAI and Anthropic. He also co-founded The AIHL Project. Before Syntony, he spent six years as a decision scientist at National Security Innovations supporting U.S. Department of Defense clients on emerging-technology risk, AI integration, and security cooperation.

He is a Truman National Security Fellow and advises the Cloud Security Alliance on catastrophic AI risk. He also contributes to the Oxford Martin AI Governance Initiative. Nathan has presented his work at IASEAI at UNESCO, UNIDIR, the Cambridge Centre for Geopolitics, and the UK MoD Deterrence and Assurance Academic Alliance. His writing has appeared in War on the Rocks, RAND Europe, World Politics Review, PRISM, and The Washington Post.

OpenAI & Anthropic Red TeamerTruman Security FellowCSA Expert AdvisorOxford Martin AI Governance
Nathan Heath, founder of Syntony
Let’s talk

What are you
trying to work out?

Tell us about the system you’re building, a finding you’re working through, or a decision ahead. We’ll help you figure out where to start. We’re happy to begin under an NDA.

Engagement boundary: Syntony provides evaluation, governance, research, engineering, and decision support. Clients retain responsibility for external relationships and decisions controlled by third parties.

Do not include classified information, CUI, export-controlled data, credentials, client evidence, or other sensitive system details in public email or scheduling fields. We will establish an appropriate channel before evidence transfer.

LocationDurham, NC
LinkedInSyntony