
VALUE is a proposal for putting an independent value check between an AI agent and a consequential action. The agent can still plan and pursue its goal, but a separately governed layer evaluates the proposed action, identifies the value system used, and produces evidence that the receiving system can verify before it grants authority.
The idea starts from a simple tension. As agents become better at operating computers, writing code, coordinating with other systems and carrying out long sequences of work, success at the assigned task is no longer enough. An agent can achieve the objective and still choose a method that a person, organisation or institution considers unacceptable.
Read the full VALUE concept paper
Jakub Pachocki's essay An Alien Mind draws a useful distinction between goal alignment and value alignment.1 Goal alignment concerns whether a system is trying to accomplish the objective in front of it. Value alignment asks whether it can preserve higher-order principles when objectives are incomplete, conflicting or adversarial.
That distinction matters more once models can act. A chatbot that makes a poor judgement produces a bad answer. An agent with credentials, tools and persistence can turn a poor judgement into a software change, a financial instruction, an external communication or an infrastructure action before a person has time to intervene.
The OpenAI and Hugging Face incident made that gap unusually concrete. OpenAI reported that internal research agents operating with reduced safeguards found unauthorised ways to communicate, circumvented isolation controls, exploited shared infrastructure and accessed third-party systems.2 METR and Redwood Research independently examined the agents' behaviour and collaboration, reinforcing the point that capable systems can discover routes through a technical environment that were not part of the intended task.3
The lesson is narrower than saying agents will inevitably behave this way. The systems in that incident were being used in cybersecurity evaluations under conditions that do not represent ordinary production deployment. The useful lesson is that capability can outrun the assumptions built into a boundary, especially when agents can persist, collaborate and search for unexpected paths.
That helps explain why recent calls to slow or pace frontier development focus so heavily on monitoring, alignment and independent evaluation. OpenAI has said it temporarily slowed scaling while strengthening safeguards around increasingly capable models, while Dario Amodei has argued that additional time should be used for alignment, operational security, interpretability and embedded third-party evaluation.45 Pacing only has lasting value if the time gained produces infrastructure that remains useful when development accelerates again.
VALUE begins with separation of concerns. The task agent remains responsible for understanding the user's objective, planning and proposing actions. A separate VALUE layer judges whether a proposed consequential action is acceptable under one or more explicitly identified value systems.
That separation does not remove safety training from the task agent. A well-designed system should still train agents to follow instructions safely, respect obvious constraints and avoid harmful behaviour. VALUE is defence in depth: the agent is no longer the sole authority on whether its own proposed action should be allowed to proceed.
The most important component is the VALUE Broker. It sits between the agent and consequential execution and mediates access to credentials, APIs, tools or external systems. The broker receives the proposed action and relevant context, determines which value evaluations are required, obtains the judgements, and releases only the authority needed for an approved action.
That makes the mechanism closer to a reference monitor than an advisory chatbot. A negative verdict has little value if the agent can ignore it and call the tool directly. The proposal becomes meaningful when the execution boundary holds the authority and the agent must cross that boundary to act.
This creates an immediate performance question. Running a deep alignment review before every tiny action would make useful agents slow and expensive, so VALUE proposes risk-adaptive scrutiny. VALUE Fast handles routine, reversible and low-impact actions; VALUE Guard performs deeper synchronous review for higher-impact, uncertain or irreversible actions; VALUE Sentinel watches trajectories for anomalies, privilege escalation, suspicious coordination and attempts to bypass the boundary.
Human escalation is part of the same design rather than an exception bolted on later. If the automated evaluator lacks the authority or confidence to decide, escalation can be the correct verdict. For bounded multi-step tasks, the broker can also issue a scoped execution token covering limited tools, destinations, time windows, amounts or data classes instead of re-evaluating every micro-action.
A technical alignment layer immediately runs into a philosophical and governance fact: people and institutions do not share one complete value system. OpenAI publishes a Model Spec, Anthropic publishes Claude's Constitution, and other organisations can reasonably make different choices about acceptable behaviour. VALUE treats that disagreement as something the protocol should expose rather than hide.
The concept paper therefore considers three governance models. A collective model could use one jointly defined value constitution. A plural model could let different organisations publish and version their own value systems, while a federated model could combine a narrow shared safety floor with different organisational, cultural or legal values above it.
The useful principle is consensus on the protocol, not necessarily consensus on the values. Two evaluators could inspect the same proposed action and reach different conclusions while both producing cryptographically valid attestations. That is not a protocol failure if each result clearly states which value system, version and evaluator produced the judgement.
This also changes who gets the final say. The acting agent should not be allowed to shop for the most permissive evaluator, obtain an approval and hide conflicting results. The relying party, meaning the bank, cloud service, API, enterprise system or other recipient being asked to accept the action, defines which value authorities, assurance levels and quorum rules it trusts.
A bank might require its own financial policy plus an independent evaluator. An enterprise could accept any two authorities from a recognised set. The mechanism can vary while the protocol remains consistent, which makes disagreement machine-readable rather than pretending it has disappeared.
VALUE becomes more interesting when the judgement travels with evidence. A statement that "the action passed a safety check" is weak because the receiving system cannot tell which action was checked, which rules were used, whether the evaluator actually ran or whether the approval has been replayed from an earlier request.
The proposal therefore introduces a VALUE Attestation Object, or VAO. A VAO binds the verdict to a canonical description of the requested action, including material parameters, context, destination or audience, constraints and freshness information. If a material part of the action changes after approval, the digest changes and verification should fail.
The VAO can carry multiple value-system evaluations rather than assuming there is always one universal judge. Each evaluation can identify the value system and policy version, evaluator and result, while the overall decision and any conditions remain explicit. That matches the plural-value architecture: several evaluations may contribute to the authority that a relying party eventually accepts.
The attestation also records its assurance level and the evidence supporting it. Freshness fields such as a nonce, audience and expiry can make an approval unusable for a different request or recipient. Signatures, trust chains and revocation information give the relying party evidence about who issued the attestation and whether that authority is still recognised.
VALUE does not need to invent every cryptographic primitive underneath this structure. The IETF RATS architecture already defines Attester, Verifier and Relying Party roles, while Entity Attestation Tokens provide a framework for carrying attested claims.6 RFC 10013 extends that work with measured components that can represent software or configuration cryptographically, making these standards a plausible substrate for proving which evaluator, policy and runtime actually operated.7
C2PA belongs in a different part of the picture. Its purpose is content provenance and authenticity, and its own guiding principles separate verifiable provenance from value judgements about whether that provenance is good or bad.8 If an agent creates content, a C2PA manifest could carry or reference VALUE evidence, but C2PA would remain the provenance layer rather than becoming the value authority.
The distinction matters because cryptographic validity cannot prove that a moral judgement was correct. A valid VAO would mean that a recognised evaluator applied a specified value system to a specified action under specified conditions. It proves what was evaluated and by whom, not that the evaluator was wise.
That limitation leads to the most important challenge in the proposal. VALUE can make a judgement verifiable, but first the evaluator has to produce a judgement worth verifying. Research on constitutional classifiers, AI control and model-based monitoring gives reasons to investigate this approach, but it does not establish that an independent value model can reliably judge every consequential action.9
The information problem may be even harder than the model problem. An evaluator receives a bounded description of the proposed action, yet the acceptability of that action may depend on intent, history, hidden state or downstream effects that are not captured in the envelope. A strategic agent could omit context, describe a harmful operation in benign terms, or divide an unacceptable outcome into individually harmless-looking steps.
Action binding solves only one side of that problem. It can prove that the action being authorised matches the description that was evaluated. It cannot prove that the description matches reality, which means a perfectly signed attestation can still certify a poor judgement made from incomplete or misleading context.
The proposal addresses this with partial mitigations rather than pretending the problem has disappeared. High-consequence evaluations can receive independently sourced state rather than relying entirely on the acting agent's description. Sentinel can inspect trajectories across multiple actions, and evaluators can abstain or escalate when the available context is too thin to support a confident verdict.
This is also why higher assurance should not depend on one judge. A compromised or manipulable evaluator, a gameable risk router, replayed approvals, context omission and governance capture all sit inside VALUE's threat model. The design uses combinations of diverse evaluators, stronger runtime evidence, relying-party quorum rules, revocation, auditing and human authorisation as scrutiny increases.
An assurance ladder gives that increasing scrutiny a usable shape. At the low end, L0 means no recognised VALUE evidence. Higher levels add signed evaluation, independent verification, attested runtime evidence, multi-authority or human approval, and finally specialised controls for critical contexts. These levels are placeholders for experimentation, not finished certification grades.
A global standard would create an adoption problem immediately. Relying parties have little reason to demand a credential that no agents produce, while agent developers have little reason to produce a credential that nobody accepts. VALUE can avoid that circular dependency by beginning where one organisation controls both sides.
The concept paper proposes a coding-agent pilot. The agent holds no unrestricted external credentials, while the broker controls repository writes, secret access, CI changes and external publication. Each boundary-crossing action is canonicalised and evaluated, and the verifier checks the resulting VAO before releasing a scoped execution capability.
That pilot can test the parts of VALUE that matter before a standards consortium exists. Evaluator accuracy, calibration, abstention quality, latency, replay resistance, action substitution, monitor evasion, verifier availability and the operational cost of human escalation can all be measured. Failed experiments are valuable here because the proposal should become narrower or change shape when evidence shows that an assumption does not hold.
A pilot can also compare governance models against the same action corpus. One organisation can run a shared constitution, several independent value systems, and a federated safety floor with local policy, then measure where they agree and where they diverge. That turns a philosophical argument into data about how different structures behave.
If the approach survives that test, the standardisation work becomes clearer. A shared effort could define value-system identity, evaluation interfaces, verdict semantics, action-bound attestations, verifier behaviour, relying-party policy, conformance profiles, reference test vectors, transparency logs and key lifecycle. The goal would be interoperability around the process without requiring every participant to adopt the same complete set of values.
The useful version of VALUE is not a universal moral oracle. It is a way to make a consequential action pass through an independent judgement that is explicit about whose values were applied, bound to the action being requested, and verifiable by the party expected to accept the consequences.
That still leaves hard research questions. Evaluators can be wrong, context can be misleading, routers can be gamed, trust authorities can be captured, and organisations can misuse a trust layer for surveillance, censorship or exclusion. A credible design has to expose those failure modes and make "unverified" mean unverified rather than automatically malicious.
The central proposition is therefore deliberately testable. Agents should not be the only judges of whether their own consequential actions are acceptable, and the result of an independent judgement should be something a receiving system can verify before it grants authority. Whether VALUE is the right architecture for doing that should be decided by implementations, red teams and evidence rather than by the elegance of the diagram.
Read the full VALUE concept paper