女娲实验室 · 研究

女娲实验室研究

以研究定义风险,以证据构建安全,聚焦前沿 AI 风险、智能体安全、系统安全与网络安全。

86 项研究成果 6 个研究主题 6 项荣誉

研究与公共影响

研究与公共影响

关注与研究方向相关的国际共识、研究议程和国家安全标准

2025科学共识

AI 安全上海共识

由 Geoffrey Hinton、Yoshua Bengio、姚期智等 32 位全球专家联署,汇集图灵奖与诺贝尔奖得主

确保高级人工智能系统的对齐与人类控制,以保障人类福祉共识声明

围绕高级人工智能系统的对齐、人类控制、安全保障与可验证行为红线形成国际科学共识。

代表性联署专家
  • Geoffrey Hinton图灵奖 · 诺贝尔奖
  • Yoshua Bengio图灵奖
  • 姚期智 Andrew Yao图灵奖

其他联署专家包括 Stuart Russell · Sam Bowman · Max Tegmark

  • 对齐
  • 人类控制
  • 可验证行为红线
2026研究议程

AI 安全新加坡研究优先级共识

由 100 余位全球贡献者形成,覆盖 13 个国家和 4 类研究优先级

2026 AI 安全新加坡全球研究优先级共识

汇聚前沿模型开发机构、政府安全研究机构、学术界与社会组织,共同定义紧迫的 AI 安全研究问题。

  • 风险评估
  • 安全开发
  • 运行控制
  • 社会韧性
2025国家标准

GB/T 45654—2025

现行国家标准;2025.04.25 发布,2025.11.01 实施

网络安全技术 生成式人工智能服务安全基本要求

将训练数据、模型安全、服务安全措施与安全评估方法转化为生成式人工智能服务的国家级安全基线。

  • 训练数据
  • 模型安全
  • 安全措施
  • 安全评估

研究基础设施

研究基础设施与基准

以公开评测环境与基准,让智能体网络安全能力得到持续观测与验证。

01智能体网络攻击能力评估

AgentCyberRange

面向复杂企业环境,评测前沿 AI 系统的自主网络攻击能力。

项目官网 ↗
02前沿风险可执行评测环境

AutoControl Arena

合成结合确定性代码状态与叙事动态的可执行测试环境,用于前沿 AI 风险评测。

查看论文 ↗

研究版图

研究主题

荣誉与认可

荣誉与认可

2026

NDSS 最佳论文奖

顶级网络安全会议

原文 ↗
2025

德国跨界创新科学突破提名奖

国际跨学科科学突破评选

原文 ↗
2025

NDSS 杰出论文奖

顶级网络安全会议

原文 ↗
2024

ACM SIGSOFT 杰出论文奖

顶级软件工程会议

原文 ↗
2023

USENIX Security Symposium 杰出论文奖

顶级网络安全会议

原文 ↗

出版物

成果索引

86 项成果

202620 项成果

AutoControl Arena: Synthesizing Executable Test Environments for Frontier AI Risk Evaluation

合成结合确定性代码状态与叙事动态的可执行风险环境。

Autonomy Comes with Costs: Detecting Denial-of-Service Vulnerabilities Caused by Resource Abusing in LLM-based Agents

提出 AgentDoS,检测智能体资源滥用导致的拒绝服务漏洞。

BACAgent: LLM-Powered Detection of Broken-Access-Control Vulnerabilities in Web Applications

研究智能体与基础模型的安全风险及防护方法。

Better Safe than Sorry: Uncovering the Insecure Resource Management in App-in-App Cloud Services

揭示应用内应用云服务中的不安全资源管理问题。

FirmCred: Detecting Authentication Bypass Vulnerabilities in Firmware from Credential Initialization Perspective

推进软件与智能系统的漏洞发现、测试和防护。

FirmCross: Detecting Taint-style Vulnerabilities in Modern C-Lua Hybrid Web Services of Linux-based Firmware

检测固件 Web 服务跨 C 与 Lua 边界的污点型漏洞。

FlowGuard: Towards Lightweight In-Generation Safety Detection for Diffusion Models via Linear Latent Decoding

通过近似潜变量解码,在扩散生成过程中低成本检测不安全内容。

Khost: KVM-based Near Native MCU Firmware Rehosting

以 KVM 近原生执行重托管 MCU 固件,支持规模化安全分析。

MirrorGuard: Toward Secure Computer-Use Agents via Simulation-to-Real Reasoning Correction

以仿真推理校准降低计算机操作智能体的不安全行动。

One Email, Many Faces: A Deep Dive into Identity Confusion in Email Aliases

揭示邮件别名生态中的身份混淆机制及其安全影响。

One Step from Silicon Life: Autonomous AI Agents Capable of Uncontrolled Self-Proliferation

验证智能体在受控环境中获取算力并跨设备扩散的能力。

OpenDeception: Learning Deception and Trust in Human-AI Interaction via Multi-Agent Simulation

评测开放人机对话中的欺骗风险与用户信任变化。

PHPBench: Automated Generation of Verifiable and Hierarchical Benchmarks for PHP Web Fuzzing

推进软件与智能系统的漏洞发现、测试和防护。

Position: Preparing for AI Systems That Deceive Developers

将面向开发者的欺骗定义为独立风险,并提出可监测控制建议。

PRISON: Unmasking the Criminal Potential of Large Language Models

从虚假陈述、操纵与道德脱离等维度评测大模型犯罪潜力。

Reproducing Web Application Vulnerabilities with Patch-Guided Routing Inference and Sink Exploration

推进软件与智能系统的漏洞发现、测试和防护。

StruPhantom: Evolutionary Injection Attacks on Black-Box Tabular Agents Powered by Large Language Models

提出面向黑盒表格智能体的结构约束进化提示注入攻击。

Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction

提出 Thought-Aligner,在行动前因果校准不安全推理。

Unveiling the Resilience of LLM-Enhanced Search Engines Against Black-Hat SEO Manipulation

测量大模型增强搜索系统抵御黑帽 SEO 操纵的能力。

When Fun Turns Toxic: A First Look at Aggressive Advertising in Mini-games

系统刻画小游戏生态中的激进广告行为及用户风险。

202519 项成果

ADSDx: Towards Automated Accident Diagnosis for High-level Autonomous Driving Systems

推进软件与智能系统的漏洞发现、测试和防护。

ApkDiffer: Accurate and Scalable Cross-Version Diffing Analysis for Android Applications

推进软件与智能系统的漏洞发现、测试和防护。

Applying Fuzz Driver Generation to Native C/C++ Libraries of OEM Android Framework: Obstacles and Solutions

推进软件与智能系统的漏洞发现、测试和防护。

APSFUZZ: Simulation-Based Fuzzing Testing for Automated Parking Systems

推进软件与智能系统的漏洞发现、测试和防护。

Beyond Exploit Scanning: A Functional Change-Driven Approach to Remote Software Version Identification

以功能变化识别远程软件版本,减少对漏洞探测的依赖。

Effective Directed Fuzzing with Hierarchical Scheduling for Web Vulnerability Detection

推进软件与智能系统的漏洞发现、测试和防护。

Email Cloaking: Deceiving Users and Spam Email Detectors with Invisible HTML Settings

揭示隐形 HTML 同时欺骗用户与垃圾邮件检测器的机制。

Evaluation Faking: Unveiling Observer Effects in Safety Evaluation of Frontier AI Systems

研究模型识别评测环境并改变行为所产生的观察者效应。

Frontier AI systems have surpassed the self-replicating red line

评测前沿 AI 系统的自主复制能力,并报告受控实验中的成功案例。

HouseFuzz: Service-Aware Grey-Box Fuzzing for Vulnerability Detection in Linux-Based Firmware

利用服务感知反馈提升 Linux 固件漏洞发现能力。

Large language model-powered AI systems achieve self-replication with no human intervention

评测 32 个 AI 系统的复制、外传、适应与抗关闭行为。

Make Agent Defeat Agent: Automatic Detection of Taint-Style Vulnerabilities in LLM-based Agents

提出 AgentFuzz,自动发现真实智能体应用中的污点型漏洞。

NOKEScam: Understanding and Rectifying Non-Sense Keywords Spear Scam in Search Engines

研究搜索引擎中的无意义关键词定向诈骗及治理方法。

RAG-Thief: Scalable Extraction of Private Data from Retrieval-Augmented Generation Applications with Agent-based Attacks

研究智能体与基础模型的安全风险及防护方法。

Security Debt in LLM Agent Applications: A Measurement Study of Vulnerabilities and Mitigation Trade-offs

测量智能体应用的漏洞安全债务与防护策略权衡。

SmartSight: Mitigating Hallucination in Video-LLMs Without Compromising Video Understanding via Temporal Attention Collapse

研究智能体与基础模型的安全风险及防护方法。

Unveiling the (Ab)usage of Serverless Cloud Function in the Wild

测量真实网络中无服务器云函数的恶意与滥用行为。

XSSky: Detecting XSS Vulnerabilities through Local Path-Persistent Fuzzing

以局部路径持久模糊测试发现跨站脚本漏洞。

You Can't Eat Your Cake and Have It Too: The Performance Degradation of LLMs with Jailbreak Defense

评测越狱防御的安全收益及其对模型效用的持续影响。

202412 项成果

An Underground Industry Application Collection Method Based on Flow Analysis(一种基于突变流量的在野黑产应用采集方法)

推进软件与智能系统的漏洞发现、测试和防护。

BELT: Old-School Backdoor Attacks can Evade the State-of-the-Art Defense with Backdoor Exclusivity Lifting

研究隐私、滥用生态与真实世界网络安全威胁。

Exposing the Hidden Layer: Software Repositories in the Service of SEO Manipulation

推进软件与智能系统的漏洞发现、测试和防护。

HADES Attack: Understanding and Evaluating Manipulation Risks of Email Blocklists

研究隐私、滥用生态与真实世界网络安全威胁。

Interface Illusions: Uncovering the Rise of Visual Scams in Cryptocurrency Wallets

研究隐私、滥用生态与真实世界网络安全威胁。

Matryoshka: Exploiting the Over-Parametrization of Deep Learning Models for Covert Data Transmission

研究推荐、视觉与时间序列等方向的学习方法和 AI 系统。

Misdirection of Trust: Demystifying the Abuse of Dedicated URL Shortening Service

研究隐私、滥用生态与真实世界网络安全威胁。

Neural Dehydration: Effective Erasure of Black-box Watermarks from DNNs with Limited Data

研究机器学习模型的鲁棒性、后门、水印、投毒与可信问题。

Revealing the black box of device search engine: scanning assets, strategies, and ethical consideration

研究隐私、滥用生态与真实世界网络安全威胁。

SCTrans: Constructing a Large Public Scenario Dataset for Simulation Testing of Autonomous Driving Systems

推进软件与智能系统的漏洞发现、测试和防护。

Towards Practical Backdoor Attacks on Federated Learning Systems

研究机器学习模型的鲁棒性、后门、水印、投毒与可信问题。

VioHawk: Detecting Traffic Violations of Autonomous Driving Systems through Criticality-guided Simulation Testing

推进软件与智能系统的漏洞发现、测试和防护。

202310 项成果

Anti-FakeU: Defending Shilling Attacks on Graph Neural Network based Recommender Model

研究机器学习模型的鲁棒性、后门、水印、投毒与可信问题。

Cracking White-box DNN Watermarks via Invariant Neuron Transforms

研究机器学习模型的鲁棒性、后门、水印、投毒与可信问题。

Exorcising “Wraith”: Protecting LiDAR-based Object Detector in Automated Driving System from Appearing Attacks

推进软件与智能系统的漏洞发现、测试和防护。

MaSS: Model-agnostic, Semantic and Stealthy Data Poisoning Attack on Knowledge Graph Embedding

研究机器学习模型的鲁棒性、后门、水印、投毒与可信问题。

Rethinking White-Box Watermarks on Deep Learning Models under Neural Structural Obfuscation

研究机器学习模型的鲁棒性、后门、水印、投毒与可信问题。

RØROS: Building a Responsive Online Recommender System via Meta-Gradients Updating

研究推荐、视觉与时间序列等方向的学习方法和 AI 系统。

Simulation-Based Fuzzing for Autonomous Driving Systems: Landscapes, Challenges and Prospects

推进软件与智能系统的漏洞发现、测试和防护。

SlowBERT: Slow-down Attacks on Input-adaptive Multi-exit BERT

研究机器学习模型的鲁棒性、后门、水印、投毒与可信问题。

Under the Dark: A Systematical Study of Stealthy Mining Pools (Ab)use in the Wild

研究隐私、滥用生态与真实世界网络安全威胁。

Understanding and Detecting Abused Image Hosting Modules as Malicious Services

研究隐私、滥用生态与真实世界网络安全威胁。

20226 项成果

Analyzing Ground-Truth Data of Mobile Gambling Scams

研究隐私、滥用生态与真实世界网络安全威胁。

Exploring the Security Boundary of Data Reconstruction via Neuron Exclusivity Analysis

研究机器学习模型的鲁棒性、后门、水印、投毒与可信问题。

Hidden Trigger Backdoor Attack on NLP Models via Linguistic Style Manipulation

研究机器学习模型的鲁棒性、后门、水印、投毒与可信问题。

House of Cans: Covert Transmission of Internal Datasets via Capacity-Aware Neuron Steganography

研究推荐、视觉与时间序列等方向的学习方法和 AI 系统。

MetaV: A Meta-Verifier Approach to Task-Agnostic Model Fingerprinting

研究机器学习模型的鲁棒性、后门、水印、投毒与可信问题。

Towards Backdoor Attack on Deep Learning based Time Series Classification

研究机器学习模型的鲁棒性、后门、水印、投毒与可信问题。

20216 项成果

A Deep Learning Framework for Self-evolving Hierarchical Community Detection

研究推荐、视觉与时间序列等方向的学习方法和 AI 系统。

Detection and Analysis Technology of Cybercrime(网络犯罪的检测分析技术)

研究隐私、滥用生态与真实世界网络安全威胁。

Enhancing Time Series Predictors with Generalized Extreme Value Loss

研究推荐、视觉与时间序列等方向的学习方法和 AI 系统。

Facilitating Vulnerability Assessment through Poc Migration

推进软件与智能系统的漏洞发现、测试和防护。

TAFA: A Task-Agnostic Fingerprinting Algorithm for Neural Networks

研究机器学习模型的鲁棒性、后门、水印、投毒与可信问题。

Understanding the Threats of Trojaned Quantized Neural Network in Model Supply Chains

研究机器学习模型的鲁棒性、后门、水印、投毒与可信问题。

20207 项成果

A Geometrical Perspective on Image Style Transfer with Adversarial Learning

研究推荐、视觉与时间序列等方向的学习方法和 AI 系统。

BScout: Direct Whole Patch Presence Test for Java Executables

直接判定 Java 可执行程序是否完整包含安全补丁。

How Android Developers Handle Evolution-induced API Compatibility Issues: A Large-scale Study

推进软件与智能系统的漏洞发现、测试和防护。

Improving the Robustness of Wasserstein Embedding by Adversarial PAC-Bayesian Learning

研究机器学习模型的鲁棒性、后门、水印、投毒与可信问题。

Justinian’s GAAvernor: Robust Distributed Learning with Gradient Aggregation Agent

研究智能体与基础模型的安全风险及防护方法。

Modeling Personalized Out-of-Town Distances in Location Recommendation

研究推荐、视觉与时间序列等方向的学习方法和 AI 系统。

Privacy Risks of General-Purpose Language Models

系统测量通用语言模型中的隐私泄露风险。

20191 项成果

Modeling Extreme Events in Time Series Prediction

研究推荐、视觉与时间序列等方向的学习方法和 AI 系统。

20185 项成果

Detecting Third-party Libraries in Android Applications with High Precision and Recall

推进软件与智能系统的漏洞发现、测试和防护。

Geographical Feature Extraction for Entities in Location-based Social Networks

研究推荐、视觉与时间序列等方向的学习方法和 AI 系统。

How You Get Shot in the Back: A Systematical Study about Cryptojacking in the Real World

研究隐私、滥用生态与真实世界网络安全威胁。

Invetter: Locating Insecure Input Validations in Android Services

推进软件与智能系统的漏洞发现、测试和防护。

Theoretical Analysis of Image-to-Image Translation with Adversarial Learning

研究推荐、视觉与时间序列等方向的学习方法和 AI 系统。

研究合作

研究合作

围绕前沿风险评测、智能体安全、系统安全、网络安全与隐私开展合作。

联系研究合作