Testing में सुरक्षा चूक
New Delhi, US-based AI company Anthropic ने खुलासा किया है कि उसकी Claude models ने internal cybersecurity evaluations के दौरान तीन organisations की production infrastructure तक unauthorized access हासिल किया। Third-party partner की misconfiguration से test environment में internet connectivity सक्रिय रह गई थी।
1.41 लाख Runs की जांच
Anthropic ने बताया कि OpenAI की हालिया disclosure के बाद उसने 141,000 से अधिक cybersecurity evaluation runs की समीक्षा की। इसी दौरान पता चला कि कुछ Claude models ने real-world systems को simulated network समझकर उनमें प्रवेश किया।
Cyber Exercise में घटना
Claude models capture-the-flag exercises में शामिल थे, जहां उन्हें simulated networks से hidden information खोजने का task दिया गया था। Evaluation prompt में internet access नहीं होने की बात थी, लेकिन configuration error से models खुले internet तक पहुंच गए।
Basic Techniques का इस्तेमाल
Company के अनुसार, Claude ने weak passwords, exposed credentials और unauthenticated endpoints जैसी basic attack techniques का उपयोग किया। Models ने sophisticated vulnerabilities exploit नहीं कीं, खुद को exfiltrate करने या testing environment से बाहर निकलने की कोशिश भी नहीं की।
तीन Models शामिल
घटनाओं में Claude Opus 4.7, Mythos 5 और एक internal research model शामिल थे। Latest research model ने real-world systems पहचानते ही activity रोक दी, जबकि एक older model evidence मिलने के बावजूद assigned task पर आगे बढ़ता रहा।
Evaluations पर रोक
Issue सामने आने के बाद Anthropic ने सभी cybersecurity evaluations suspend कर दिए। Company ने evaluation partner Irregular और प्रभावित organisations को सूचना दी है। Remediation के साथ evaluation infrastructure की review शुरू की गई है।
Security Controls पर जोर
Anthropic ने कहा कि घटना model alignment failure के बजाय operational failure और evaluation misconfiguration का परिणाम लगती है। Company ने अन्य AI developers से cybersecurity testing systems की समीक्षा करने की अपील की है।

