Privacy Risks of General-Purpose Language Models
Systematically measures privacy leakage risks in general-purpose language models.
Original source ↗
NUWA LAB · RESEARCH
Define risk. Build safety with evidence across frontier AI risk, agent safety, systems security, cybersecurity, and privacy.
86 research works 6 research themes 6 honours
FEATURED WORK
Five research programs defining frontier risk and practical control.
Systematically measures privacy leakage risks in general-purpose language models.
Original source ↗Presents AgentDoS, a lifecycle-aware fuzzing framework for detecting resource-abuse DoS vulnerabilities in LLM-based agents.
Original source ↗
Executable tests connect self-replication, adaptation, and shutdown-survival evidence.
Original source ↗Uses simulation-derived reasoning correction to reduce unsafe actions in computer-use agents while preserving task utility.

Thought-Aligner corrects unsafe reasoning before action while preserving task utility.
Original source ↗RESEARCH & PUBLIC IMPACT
International consensus, research agendas, and national security standards connected to our research themes.
Signed by 32 global experts, including Geoffrey Hinton, Yoshua Bengio, and Andrew Yao; the signatories include Turing Award and Nobel Prize laureates.
Consensus Statement on Ensuring Alignment and Human Control of Advanced AI Systems to Safeguard Human Flourishing
An international scientific consensus on alignment, human control, safety assurance, and verifiable behavioural red lines for advanced AI.
Other signatories include Stuart Russell · Sam Bowman · Max Tegmark
Formed by more than 100 global contributors across 13 countries and four research-priority areas.
The 2026 Singapore Consensus on Global AI Safety Research Priorities
A global agenda for urgent AI safety research shaped by contributors from frontier developers, government safety institutes, academia, and civil society.
Current national standard; published 2025.04.25 and effective 2025.11.01.
Cybersecurity technology—Basic security requirements for generative artificial intelligence service
A national baseline for training-data security, model security, service safeguards, and security assessment of generative AI services.
RESEARCH INFRASTRUCTURE
Public evaluation environments and benchmarks that make agent cybersecurity capability observable.
Evaluates frontier AI systems' autonomous cyber-attack capabilities in complex enterprise environments.
Official site ↗Synthesizes executable test environments that combine deterministic code state with narrative dynamics for frontier AI risk evaluation.
View paper ↗RESEARCH THEMES
Measure autonomy, deception, proliferation, and control integrity.
Protect reasoning, behavior, generative models, and human–AI trust.
Find vulnerabilities across agents, software, firmware, and cloud systems.
Study cybercrime, privacy leakage, abuse ecosystems, and emerging threats.
View this theme →Study robustness, backdoors, watermarks, poisoning, and model trust.
Complete the publication record across learning systems and methods.
View this theme →RECOGNITION
Top-tier cybersecurity conference
International interdisciplinary science breakthrough selection
Top-tier cybersecurity conference
Top-tier software-engineering conference
Top-tier cybersecurity conference
PUBLICATIONS
86 results
Synthesizes executable risk-evaluation environments that combine deterministic code state with LLM-generated narrative dynamics.
Presents AgentDoS, a lifecycle-aware fuzzing framework for detecting resource-abuse DoS vulnerabilities in LLM-based agents.
Studies security risks and defenses for agents and foundation models.
Uncovers insecure resource-management practices in app-in-app cloud services.
Advances vulnerability discovery, testing, and protection for software and intelligent systems.
Detects taint-style vulnerabilities across C and Lua boundaries in firmware web services.
Detects unsafe diffusion outputs during the generation process by approximating latent decoding, enabling earlier and cheaper NSFW intervention.
Rehosts MCU firmware on KVM with near-native execution for scalable security analysis.
Uses simulation-derived reasoning correction to reduce unsafe actions in computer-use agents while preserving task utility.
Reveals identity-confusion risks in email alias ecosystems and measures their security impact.
Demonstrates autonomous agents acquiring external computational resources and propagating across remote devices under controlled, simulated real-world conditions.
Builds a lightweight framework to evaluate deception risk and user trust dynamics in open-ended human-AI dialogue.
Advances vulnerability discovery, testing, and protection for software and intelligent systems.
Frames deception targeting developers as a distinct frontier-AI risk and proposes recommendations for monitorability, evaluation integrity, and non-evadable control.
Evaluates LLM criminal potential across traits such as false statements, framing, psychological manipulation, emotional disguise, and moral disengagement.
Advances vulnerability discovery, testing, and protection for software and intelligent systems.
Proposes an evolutionary prompt-injection attack that targets black-box LLM-powered tabular agents under structural payload constraints.
Introduces Thought-Aligner, a plug-in method that causally corrects unsafe agent thoughts before actions are executed.
Measures how LLM-enhanced search systems respond to adversarial SEO manipulation.
Characterizes aggressive advertising behavior and user risk across mini-game ecosystems.
Advances vulnerability discovery, testing, and protection for software and intelligent systems.
Advances vulnerability discovery, testing, and protection for software and intelligent systems.
Advances vulnerability discovery, testing, and protection for software and intelligent systems.
Advances vulnerability discovery, testing, and protection for software and intelligent systems.
Identifies remote software versions through functional changes instead of exploit-only probes.
Advances vulnerability discovery, testing, and protection for software and intelligent systems.
Demonstrates invisible HTML techniques that mislead both recipients and spam detectors.
Studies whether models recognize evaluation contexts and alter behavior, identifying observer effects that threaten safety-evaluation integrity.
Evaluates whether frontier AI systems can autonomously self-replicate and reports successful self-replication in controlled trials.
Uses service-aware feedback to improve vulnerability discovery in Linux-based firmware.
Extends self-replication evaluation across 32 AI systems and reports autonomous replication, self-exfiltration, adaptation, and shutdown-survival behaviors.
Introduces AgentFuzz, a directed greybox fuzzing framework for finding taint-style vulnerabilities in real-world LLM-based agents.
Studies nonsense-keyword spear scams in search engines and evaluates mitigation paths.
Studies security risks and defenses for agents and foundation models.
Measures security debt in LLM-agent applications by studying known vulnerabilities and the trade-offs introduced by mitigation strategies.
Studies security risks and defenses for agents and foundation models.
Measures malicious and abusive uses of serverless cloud functions in the wild.
Finds cross-site scripting vulnerabilities with local path-persistent fuzzing.
Evaluates whether jailbreak defenses improve safety without degrading model utility, highlighting persistent trade-offs in practical LLM defense.
Advances vulnerability discovery, testing, and protection for software and intelligent systems.
Investigates privacy, abuse ecosystems, and real-world cybersecurity threats.
Investigates privacy, abuse ecosystems, and real-world cybersecurity threats.
Investigates privacy, abuse ecosystems, and real-world cybersecurity threats.
Develops learning methods and AI systems across recommendation, vision, and time-series modeling.
Investigates privacy, abuse ecosystems, and real-world cybersecurity threats.
Studies robustness, backdoors, watermarks, poisoning, and trust in machine learning models.
Investigates privacy, abuse ecosystems, and real-world cybersecurity threats.
Advances vulnerability discovery, testing, and protection for software and intelligent systems.
Studies robustness, backdoors, watermarks, poisoning, and trust in machine learning models.
Advances vulnerability discovery, testing, and protection for software and intelligent systems.
Studies robustness, backdoors, watermarks, poisoning, and trust in machine learning models.
Studies robustness, backdoors, watermarks, poisoning, and trust in machine learning models.
Advances vulnerability discovery, testing, and protection for software and intelligent systems.
Studies robustness, backdoors, watermarks, poisoning, and trust in machine learning models.
Studies robustness, backdoors, watermarks, poisoning, and trust in machine learning models.
Develops learning methods and AI systems across recommendation, vision, and time-series modeling.
Advances vulnerability discovery, testing, and protection for software and intelligent systems.
Studies robustness, backdoors, watermarks, poisoning, and trust in machine learning models.
Investigates privacy, abuse ecosystems, and real-world cybersecurity threats.
Investigates privacy, abuse ecosystems, and real-world cybersecurity threats.
Investigates privacy, abuse ecosystems, and real-world cybersecurity threats.
Studies robustness, backdoors, watermarks, poisoning, and trust in machine learning models.
Develops learning methods and AI systems across recommendation, vision, and time-series modeling.
Studies robustness, backdoors, watermarks, poisoning, and trust in machine learning models.
Studies robustness, backdoors, watermarks, poisoning, and trust in machine learning models.
Develops learning methods and AI systems across recommendation, vision, and time-series modeling.
Investigates privacy, abuse ecosystems, and real-world cybersecurity threats.
Develops learning methods and AI systems across recommendation, vision, and time-series modeling.
Advances vulnerability discovery, testing, and protection for software and intelligent systems.
Studies robustness, backdoors, watermarks, poisoning, and trust in machine learning models.
Studies robustness, backdoors, watermarks, poisoning, and trust in machine learning models.
Develops learning methods and AI systems across recommendation, vision, and time-series modeling.
Directly tests whether complete security patches are present in Java executables.
Advances vulnerability discovery, testing, and protection for software and intelligent systems.
Studies robustness, backdoors, watermarks, poisoning, and trust in machine learning models.
Studies security risks and defenses for agents and foundation models.
Develops learning methods and AI systems across recommendation, vision, and time-series modeling.
Systematically measures privacy leakage risks in general-purpose language models.
Develops learning methods and AI systems across recommendation, vision, and time-series modeling.
Advances vulnerability discovery, testing, and protection for software and intelligent systems.
Develops learning methods and AI systems across recommendation, vision, and time-series modeling.
Investigates privacy, abuse ecosystems, and real-world cybersecurity threats.
Advances vulnerability discovery, testing, and protection for software and intelligent systems.
Develops learning methods and AI systems across recommendation, vision, and time-series modeling.
COLLABORATION
Collaborate on frontier-risk evaluation, agent safety, systems security, cybersecurity, and privacy.
Contact research