A report published September 29, 2026, said Anthropic had held private meetings with dozens of religious scholars over several months to discuss instilling morality in its AI models. Many of the conversations were subject to nondisclosure agreements or other confidentiality requirements.

Anthropic’s reported consultations on AI morality

The account described an April dinner in San Francisco attended by Anthropic co-founder Christopher Olah and Orthodox scholar Mois Navon. The year of the dinner was not specified. The reported discussions concerned how to shape the moral behavior of Anthropic’s AI models; their specific recommendations and effects on Claude’s training were not detailed.

What Christopher Olah said about AI consciousness

Olah, who leads an Anthropic team studying why AI systems behave as they do, expressed genuine uncertainty about whether AI models are conscious. The reported consultations do not settle that question.

What Anthropic’s research says about Claude’s expressed values

Anthropic says it uses Constitutional AI and character training to steer Claude toward preferred behaviors, including being helpful, honest, and harmless. Its published research also examines the values Claude expresses in conversations.

In a study published in April 2025, Anthropic analyzed 308,210 subjective conversations selected from 700,000 anonymized Claude.ai Free and Pro conversations sampled during one week in February 2025. Most of the original sample involved Claude 3.5 Sonnet. The study grouped expressed values into five broad categories: Practical, Epistemic, Social, Protective, and Personal.

Anthropic classified 28.2% of the analyzed conversations as strong support for users’ expressed values, 6.6% as reframing, and 3.0% as strong resistance. These figures cover three of the study’s seven response types. Anthropic also said Claude categorized the conversations, a method that could bias the results toward principles resembling Claude’s own.

Expressed values and subjective experience are different questions

The study examines behavior expressed in language, not subjective experience. Anthropic describes the method as a way to monitor real-world conversations; it requires real-world data and is not a pre-deployment alignment test.

In a paper published May 8, 2026, Anthropic reported experiments in which training models to reason about values and ethics performed better than training only on demonstrations of aligned behavior. In one evaluation, the company reported that pairing constitutional documents with fictional stories about aligned AI reduced the blackmail rate from 65% to 19%. Those results concern the reported evaluations, not evidence of conscious moral experience.