Exploring Anthropic's groundbreaking 'Constitutional Classifiers' research that withstood 3,000+ hours of jailbreak attempts with a $15,000 bounty, using separate classifier models as effective AI safety guardrails.
Лучшая цитата из Unbreakable AI Guardrails
“
The key innovation here is that instead of trying to make the main AI model refuse harmful requests, they're using separate 'classifier' models that act as guardrails. These classifiers are trained using what they call a 'constitution' - basically natural language rules defining what's allowed and what's not.
”
S
Generated by Song
Вопрос для ввода
Help me find this paper Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
Голоса ведущих
Lena
Miles
Источники знаний
Создано выпускниками Колумбийского университета в Сан-Франциско
BeFreed объединяет глобальное сообщество из 1,000,000 любознательных умов
AI models are fast but unpredictable. Learn how harness engineering creates the safety systems needed to turn raw AI power into reliable production code.
Refusal training isn't enough when dangerous knowledge remains inside AI. Learn how selective forgetting can secure frontier models without breaking them.