AI

Anthropic’s Claude 4 safety tech and the EU AI Act deadline shift

Anthropic's updated Constitutional Classifiers reduce jailbreak success rates from 86% to 4.4%. This safety technology becomes critical as the EU AI Act's Article 50 transparency obligations remain set for August 2, 2026, despite extensions for high-risk systems.

Anthropic's Claude 4 safety tech and the EU AI Act deadline shift

The EU compliance clock keeps ticking

The Council of the European Union gave final approval to the Digital Omnibus simplification package on June 29, 2026, pushing the compliance deadline for stand-alone high-risk AI systems under Annex III from August 2, 2026, to December 2, 2027. This reprieve does not touch the Article 50 transparency obligations. Those requirements, which mandate disclosing AI interactions, remain on the original August 2, 2026, schedule. You should not ignore these rules just because the high-risk deadline moved. I consider this a distinction that companies must grasp to avoid penalties. High-risk AI systems embedded in regulated products, such as medical devices and machinery, receive a parallel 12-month extension, moving their deadline from August 2, 2027, to August 2, 2028. The high-risk tier, governed by Annex III, encompasses use cases like employment and credit scoring. This includes AI used for creditworthiness evaluation and risk assessment for life or health insurance. Providers face penalties of up to €15 million or 3 percent of global annual turnover. The extension exists because technical specifications for meeting Annex III were not ready, making the legislative text’s arrival dependent on standardization work. The Council’s June 29, 2026, action was the final procedural step, and the legislative text should appear in the EU’s Official Journal shortly.

EU AI Act Requirement Original Deadline New Deadline
Annex III Stand-alone August 2, 2026 December 2, 2027
Annex I Product-embedded August 2, 2027 August 2, 2028
Article 50 Transparency August 2, 2026 August 2, 2026
Watermarking (deployed systems) August 2, 2026 December 2, 2026
Article 5 Intimate Imagery February 2, 2025 December 2, 2026

Anthropic’s classifier performance

Anthropic’s updated Constitutional Classifiers reduce jailbreak success rates from 86% to 4.4%. The system uses an ensemble defense consisting of a linear probe and a probe-classifier ensemble. This two-stage architecture allows a lightweight first-stage classifier to screen all traffic, which then escalates suspicious exchanges to a more powerful second-stage classifier. This method limits the increase in refusal rates for harmless queries to just 0.38%. I find the minimal compute overhead of ~1% for Opus 4.0 traffic quite useful. The classifiers use synthetic data to train input and output classifiers that follow a constitution of principles defining allowed and disallowed content. For example, recipes for mustard are allowed, but recipes for mustard gas are not. The second-stage classifier screens both sides of a conversation, which makes it better able to recognize jailbreaking attempts. The updated version achieved similar effectiveness on synthetic evaluations with only a 0.38% increase in refusal rates and moderate additional compute costs. During automated evaluations, the system tested 10,000 jailbreaking prompts. In human red teaming with 183 participants, no participant successfully coerced the model to answer all ten forbidden queries with a single jailbreak.

Safety tools and regulatory timelines

The delay for high-risk systems does not cancel out transparency needs. Article 50 requires informing individuals they interact with AI. The deadline for labeling AI-generated content for systems already in deployment moves to December 2, 2026. A new prohibition on AI generating non-consensual intimate imagery also takes effect on December 2, 2026. Companies should use this time to integrate tools like Anthropic’s classifiers. How will regulators evaluate the effectiveness of these automated safety layers? Most organizations still lack a basic inventory of the AI systems they operate, as CSA research from March 2026 found, while a 2023 analysis found that 40 percent of 106 enterprise AI systems could not be cleanly classified against the Act’s risk tiers. Deployers of high-risk systems must assign human oversight and retain logs for a minimum of six months. Providers face the most demanding compliance burden, which includes conformity assessments and technical documentation. If a deployer makes a substantial modification to a high-risk system, they become a provider; for example, fine-tuning a vendor credit model on proprietary data can trigger this reclassification. The EU AI Office also gains expanded investigatory and enforcement authority, including on-site inspection powers and the ability to secure binding commitments from providers.