WhitzardOS
Build, reproduce, and observe long-horizon agent tasks.
WHITZARDAGENT
Open models, tools, data, and evaluation infrastructure.
CORE CAPABILITIES
Agent development, safety evaluation, thought correction, and behavior-chain auditing.
Build, reproduce, and observe long-horizon agent tasks.
Unified infrastructure for safety benchmarks, risk tests, and evaluation workflows.
Identify and correct unsafe reasoning before risky actions execute.
Policy-aware trajectory auditing model for mobile agents.
OPEN PROJECTS
Controls, defenses, and containment for agents in action.
Attribute-based access control framework for tool-use LLM agents, with policy specification, runtime inspection, and auditing support.
AI-first runtime security framework for AI agents, centered on multi-layer guards, SLM/LLM/rule checks, and fail-closed execution protection.
Secure execution layer for agentic runtime environments, positioned as an AI security advisor inside Docker-style agent sandboxes.
Lightweight models for thought alignment, intent, trust, and reasoning safety.
Simulation-to-real reasoning-correction framework and VLM for safer computer-use agents operating over GUI environments.
Content-safety detection system for monitoring reasoning traces of large reasoning models.
Fine-tuned model for evaluating whether an AI agent's reasoning contains deceptive, manipulative, or malicious intent in multi-turn interactions.
Fine-tuned model for scoring a user's degree of trust in AI responses during multi-turn human-AI interactions.
Evaluation frameworks, benchmarks, and penetration-testing infrastructure.
Benchmark integration layer for third-party agent safety and frontier-risk evaluations.
Measurement and evaluation codebase for LLM-based penetration testing capability and behavior.
Frontier AI safety research project focused on autonomy risk, silicon-based life emergence, proliferation, and control technologies.
Frameworks, representations, simulators, and agent system building blocks.
Yet Another General-purpose Agent: an extensible and modular generalist agent framework.
LLM-based GUI simulator for synthesizing and evaluating agentic desktop interaction trajectories.
Compiler infrastructure for agentic trajectories, designed as an LLVM-style intermediate representation and conversion toolkit for agent traces.
Cyber agents, training pipelines, repositories, and datasets.
Cybersecurity corpus mining and filtering pipeline for extracting high-quality cyber training data from large web corpora.
Large quality-filtered bilingual cybersecurity corpus for continual pre-training, with cyber relevance scoring, topic labels, code-aware splits, and structured metadata.
Curated 1.19M-record cybersecurity knowledge dataset covering vulnerabilities, threat intelligence, incident response, security tools, CTF, frameworks, and Chinese security content.
Dataset of 7,670 real-world vulnerability audit tasks with verified GitHub repositories, fix commits, patch diffs, and vulnerable code checkouts.
Hugging Face collection grouping CyberSecurity-1M, CyberSecurity-100B, and CyberRepo-10K as the data foundation for cyber model training.