AI 위협에 대해 증언하는 탑 AI 경영진
위험 평가 방법론:
앤트로픽은 생물학적 보안 및 사이버 보안 위협을 정량화하기 위해 상세한 프로세스를 갖춘 책임 있는 확장 정책을 사용하며, 특히 제어 상실 및 오정렬 위험에 중점을 두고 결과를 공개 위험 보고서에 발표합니다.
구글의 프론티어 안전 프레임워크는 업계 벤치마크를 사용하여 모델의 핵심 역량 수준을 모니터링하고, 모델이 사전 정의된 위험 역량 임계값을 초과할 경우 안전 완화 조치를 자동으로 실행합니다.
메타는 배포 전 평가, 적대적 테스트, 그리고 초지능 확장 프레임워크 및 준비 보고서에 명시된 안전 장치를 결합한 다층적 안전 접근 방식을 구현하고, 위험 완화 후에만 배포합니다.
근본적인 과제:
학계 전문가들에 따르면, AI 기업들은 확률 계산을 위한 참조 클래스가 부족하여 치명적인 위험을 정량화하는 데 근본적인 어려움을 겪고 있으며, 따라서 정량적 위험 주장을 할 때는 겸손한 자세를 유지해야 합니다.
산업 전략:
앤트로픽은 안전 프로세스를 위한 충분한 시간을 확보하기 위해 AI 개발의 속도를 조절할 것을 주장하며, 책임 있는 확장 정책은 신중한 개발 접근 방식을 유지하는 데 효과적인 것으로 입증되었습니다.
Risk Assessment Methodologies
-
Anthropic uses a responsible scaling policy with detailed processes to quantify biological security and cybersecurity threats, focusing specifically on loss of control and misalignment risks while publishing findings in public risk reports.
-
Google’s frontier safety framework monitors models against critical capability levels using industry benchmarks, automatically triggering safety mitigations when models cross predefined dangerous capability thresholds.
-
Meta implements a multi-layered safety approach combining pre-deployment evaluations, adversarial testing, and safeguards outlined in their superintelligence scaling framework and preparedness reports, deploying only after risk mitigation.
Fundamental Challenges
-
AI companies face fundamental difficulty quantifying catastrophic risks due to absence of reference classes for probability calculations, requiring humility when making quantitative risk claims according to academic experts.
Industry Strategy
-
Anthropic advocates pacing the frontier of AI development to extend time available for safety processes, with responsible scaling policies proving effective for maintaining cautious development approaches thus far.