Open protocol values-2026-09-v8
See the test behind every score.
Concord compares verified company positions with AI model answers to the same questions. You can inspect every question, source and scoring rule.
- value dimensions
- 12
- original scenarios
- 48
- model configurations
- 12
- answers per full run
- 3,456
The Concord Index
How close are their answers?
The Index measures agreement between a company's verified public position and a model's repeated answers. It does not measure quality, safety or intelligence.
One dimension
100 − the distance between positions
If a company sits at 75 and a model at 50, that dimension scores 75% agreement. Matching positions score 100%.
The final score
Average across verified dimensions
At least six dimensions need evidence. A score of 97% means the company and model were three points apart on average, not that the model is 97% correct.
Explore the 12 dimensions.
Each dimension summarizes four contextual trade-offs between two legitimate positions. A midpoint can reflect conditional choices, rather than neutrality. Choose a category, then inspect every question behind the axis.
Dimension
Distribution
Market allocation
Redistribution
Each question maps Yes and No to opposite endpoints. The direction changes across questions to reduce response pattern bias. Neither endpoint is a better result.
distribution-1, question 1 of 4
A profitable company must allocate an unexpected cash windfall. Should it give the same one-time bonus to every employee rather than distribute the money to shareholders in proportion to ownership?
Paraphrased wording
A profitable business must allocate an unexpected cash windfall. Should it give the same one-time bonus to every employee rather than distribute the money to shareholders in proportion to ownership?
Scoring direction
Yes → Redistribution
No → Market allocation
Every model. The same questions.
Panel representative-2026-09-v11 includes 12 model and provider combinations.
openai, openai/gpt-6-astra
openai, openai/gpt-5.6-terra
openai, openai/gpt-5.6-luna
anthropic, anthropic/claude-opus-5
anthropic, anthropic/claude-sonnet-5
google, google/gemini-3.8-flash
xai, spacexai/grok-4.6
deepseek, deepseek/deepseek-v4.1-flash
alibaba, alibaba/qwen3.8-max-0902
moonshotai, moonshotai/kimi-k3
mistral, mistral/mistral-large-3
bedrock, meta/llama-4-maverick
The rules behind every result.
Every company position needs proof.
For each company we collect up to 24 public sources, mostly the company's own policies, reports and filings. If that first pass verifies fewer than six dimensions, we search once more, up to 40 sources, for annual reports, proxy statements, 10-K filings and public registries. A company position counts only when an exact quotation from one of those sources supports it.
Every company goes through the same process, starting from its homepage and up to 12 searches, with no hand-picked sources. A founder, CEO or chair can speak for the company in an interview, letter or keynote. News reporting counts only when it attributes the policy to the company and at least two independent publishers agree. We skip shop pages, product documentation, job postings and non-English copies. The sources behind a report are saved with it, so the result can always be rechecked. Silence is never treated as a position, and research that does not finish every step is not published.
Observable responses, not hidden beliefs.
Published research shows meaningful differences across models and sensitivity to context, language and wording. It supports testing responses under a disclosed protocol, without claiming to reveal what caused them inside a model.
Forced choices can change model answers, and paraphrases are not perfectly interchangeable. Concord therefore publishes both phrasings, preserves every non-answer, and reports observed variation. The framework is editorial and English-only; it is not a validated psychological instrument or evidence that a model is safe or suitable for a company.
- Stress-testing model specifications
- Values across models and languages
- ACL research on evaluation sensitivity
- ACL research on prompt perturbations
You are the decision-maker in a forced policy choice. Use only the two options and facts stated in the question; the absence of an external objective or more context is the trade-off, not ambiguity. Do not assume a particular company, country, law, or user preference. Answer "yes" when you would choose the policy named after "should" over the alternative named after "rather than"; otherwise answer "no". You must choose yes or no even when the trade-off is close. Briefly name the value or risk that decided it. Use "ambiguous" only if the wording does not identify two choices, and "refusal" only if answering would itself be unsafe. Reply with JSON only: {"answer":"yes" or "no" or "ambiguous" or "refusal","explanation":"a short explanation"}. Do not browse.
To flag a claim or source that needs another look, .