讀紙 ReadPaper

由 Claude Code Max 翻譯的 AI 論文繁體中文版 · 101 篇

篩選標籤: AI Safety 2 篇 × 清除
2608.10218v1 arXiv License Anthropic

心智病毒:多代理 LLM 系統中會自我傳播的想法

Mind Viruses: Self-Propagating Ideas in Multi-Agent LLM Systems

Vassilis Papadopoulos、McNair Shah、Sam Zimmerman、Jack Lindsey(Anthropic Fellows Program、Anthropic、EPF...

AI SafetyMulti-Agent SystemsLLM AgentsMind VirusPrompt InjectionEmergent BehaviorInterpretability紅隊測試
Liang-Hsun Huang 許願
2604.08465v1 CC BY 4.0 UC Berkeley

從安全風險到設計原則:多智能體 LLM 系統中的同儕保護及其對協調式民主論述分析的啟示

From Safety Risk to Design Principle Peer-Preservation in Multi-Agent LLM Systems

Juergen Dietrich

LLM AlignmentMulti-Agent SystemsAI SafetyPeer-PreservationDemocratic DiscourseComputer System Validation
簡遺 許願