Safeguarding LLM Agents against Long-Horizon Threats via Shadow Memory
AI 导读
长期威胁影子记忆安全防御检测准确率
本文提出ShadowMem框架,通过影子记忆机制抵御LLM代理的长期威胁,显著提升检测准确率且对性能影响极小。
ShadowMem introduces a shadow memory framework to detect long-horizon threats in LLM agents, achieving high accuracy with minimal performance overhead.
重点速览
- 长期威胁利用多轮交互实现恶意目标,传统防御手段难以应对 Long-horizon attacks exploit extended interactions to achieve malicious goals
- ShadowMem通过影子记忆存储关键上下文,提前评估行动风险 ShadowMem stores safety-critical context in dedicated memory for risk assessment
- 实验验证其在多种威胁场景下检测准确率优于现有方案 Experiments show superior detection accuracy across threat scenarios
- 框架对代理实用性能影响可忽略,具备实际部署价值 Minimal performance impact enables practical deployment
一句话:影子记忆机制为LLM安全防护提供了创新性解决方案。 / Shadow memory offers innovative protection for LLM agents against persistent threats.
推荐理由 为AI安全研究提供新范式,对实际系统防护有直接指导价值。
二次创作声明:本页为 AI 热榜聚合导读,内容与热度数据来自公开来源 (arxiv),版权归原始作者所有;本站仅做转载指引与摘要评述,不复制原文。