Selected work — Ainsley Woo

Projects and research by Ainsley Woo across custom AI agents, LLM and agent security, AI red-teaming, post-quantum cryptography and malware analysis.

AgentGuard

Agent security · 2026Dual-stream anomaly detection for LLM-agent servers.

Fuses OS telemetry and LLM action logs through a Mamba + Transformer cross-attention architecture to catch prompt injection, exfiltration, resource abuse, and persistence attacks invisible to either signal alone — F1 0.95 / AUROC 0.995, beating five baselines.

SentinelLM

LLM security · 2026A runtime adversarial firewall for LLM applications.

A taxonomy-aware safety layer for LLM apps: it screens prompts through multi-layer intent analysis and is hardened with adversarial self-play, shipping a core library plus benchmark suites that score it ~500% over a Qwen baseline.

origin

Agentic coding harness · 2026A Rust-native agentic coding harness — CLI + supervised daemon.

A CLI plus a supervised daemon that run LLM-driven coding sessions locally, in Rust. Carries CI, codecov, and an OpenSSF Scorecard badge under Apache-2.0 — built for safe, observable autonomous coding.

AISOCA

AI SOC analyst · 2026An AI-augmented SOC analyst that triages and explains.

An AI-powered Security Operations Center system that triages alerts, suggests investigations, explains its reasoning, and learns from analyst corrections without model drift — multi-dimensional alert classification end to end.

QASP

Post-quantum networking · 2026A post-quantum secure protocol for agent-to-agent comms.

A post-quantum secure network protocol for AI agent-to-agent communication on TCP/QUIC with a sans-I/O state-machine core: a custom handshake combining ML-KEM-768 / ML-DSA-65 with hybrid X25519, stream multiplexing, OCSP-style revocation, capability tokens, Bayesian trust scoring, and MCP/A2A bridges.

Malware Analysis Sandbox

Malware analysis · 2025Static + dynamic malware analysis across platforms.

An enterprise-grade sandbox for static and dynamic analysis: PE/ELF/script/document parsing, YARA, entropy + packer detection, isolated Docker execution with API/syscall tracing, PCAP network analysis with IOC extraction, and Volatility-3 memory forensics.

AI Attack Validation Suite

AI red-teaming · 2026Runnable coverage of 92 AI/LLM attack types.

A red-teaming harness implementing all 92 attack types from a taxonomy of AI/LLM attacks — one runnable script per attack, pointed at OpenAI, Anthropic, or Gemini models through a unified client.

BenchAudit

Eval integrity · 2026A meta-evaluator that audits agentic-AI benchmarks.

Statically detects, then live-demonstrates, substrate-layer exploits that inflate scores on real eval harnesses (SWE-bench Verified, Terminal-Bench) — scoring each harness per the SEALED whitepaper and reporting in SARIF + Markdown.

HELIX

LLM-agent memory · 2026A unified memory architecture for LLM agents.

A layered memory substrate for LLM agents built on a unified causal-provenance graph that every layer queries, with an episodic layer and a retrieval-fuse stage — so an agent's memory is one coherent, auditable store rather than scattered caches.

JANUS

Theory-of-mind / belief · 2026A recursive belief ledger for tool-composing agents.

A reference implementation of the JANUS recursive belief ledger and PVU runtime — modelling an agent's nested beliefs about other agents (theory of mind) so tool-composing multi-agent systems reason about who knows what.

VAC

Multi-agent trust · 2026A Verifiable Agent Compact for multi-agent AI.

A trust architecture for multi-agent AI — the Verifiable Agent Compact — starting at Layer 0 with cryptographic identity and attestation so agents can prove who they are and what they ran before any cross-agent action is trusted.

CALYPSO

World models / planning · 2026A goal-projected one-step latent world model.

Component 1 of CALYPSO: a trained goal-projected one-step latent operator T̂ with a calibrated uncertainty head, plus the data/training/eval pipelines that produce it — the thesis that a one-step world model is enough for planning.

RLP

Latent planning · 2026A plan-space encoder for retrocausal latent planning.

Retrocausal Latent Planning — a plan-space encoder that represents whole plans in a latent space, letting a planner reason backward from goals rather than only forward from the current state.

AINSLEY

Efficient ML inference · 2026Runs a frontier MoE on a 16 GB laptop GPU.

Adaptive Indexer-driven N-tier Storage with Layered Entropy Yield: puts a frontier hybrid-sparse-attention MoE model (DeepSeek-V4-Flash class) onto a 16 GB laptop GPU at interactive throughput by exploiting the workload's redundancy across model architecture, per-vector quantization, and serving systems at once.

Beyond the Reflexion Ceiling

Agent training · 2026A production training recipe past the reflexion ceiling.

The π_a training recipe for the AHV-CRC component — a production recipe aimed at pushing reasoning agents past the ceiling that pure reflexion-style self-correction hits.

Predictive Vulnerability Engine

Post-quantum crypto · 2026Forecasts quantum threats to cryptographic systems.

An enterprise-grade predictive vulnerability engine (Rust + Python) combining symbolic reasoning and ML to forecast quantum threats, scan codebases across six languages for at-risk primitives, and automate FIPS-140-2 / FedRAMP-compliant migration to PQC (Kyber, Dilithium, SPHINCS+).

Zero-Trust Attack-Path Mapper

Zero-trust security · 2026Maps attack paths across a zero-trust estate.

A Rust tool that models a zero-trust environment and maps the reachable attack paths through it — surfacing how an adversary could chain identities and resources despite per-request verification.

Autonomous Malware Attribution

Malware analysis · 2026Behavioural-GNN malware attribution at scale.

An autonomous malware attribution platform in Rust + Python combining behavioural graph neural networks with rule-based detection across Windows/Linux/macOS, scaling to 10,000+ hosts with sub-second attribution and automatic MITRE ATT&CK mapping.

Continuous Supply-Chain Trust Verifier

Supply-chain security · 2026Continuously vets dependencies across ecosystems.

A Rust supply-chain platform (11-crate workspace, REST/GraphQL API, Leptos dashboard, Postgres worker queue) that continuously scans npm, PyPI, Maven, Cargo and more for typosquatting, tampering, and provenance attacks — emitting CycloneDX/SPDX SBOMs and SARIF.

NMAT

Network analysis · 2025A premium network monitor + packet analyzer.

Network Monitor & Analysis Tool — an Electron packet-analysis app with a glassmorphic, Apple-inspired UI: live capture, protocol analysis, and a frameless, blur-backdropped interface that makes deep network inspection feel premium.

Hallucinations at the Firewall

AI safety · AAAI 2026 · 2026Why AI hallucinates, and what it costs enterprises.

Presented at the AAAI 2026 Undergraduate Consortium: research on why AI systems hallucinate and how that harms enterprises and consumers — framing hallucination as a security-relevant failure mode at the application boundary.

Prompt Refiner

Web app · prompt engineering · 2026Improve prompts via best-practice refinement.

A web app that evaluates a prompt against eight prompt-engineering criteria, asks multiple-choice questions to fix the unmet ones, then generates responses from the original and refined prompts side by side.

Agendaler

Proptech · SG real estate · 2026AI-powered Singapore property search.

An AI-powered property-search platform for Singapore real estate — turning natural-language intent into ranked, explained listings.

Sports Month 2026

Event platform · 2026An enterprise event platform for Sports Month.

A full event-management platform for Sports Month 2026 — Next.js 15 + PostgreSQL 16 + TypeScript, with enterprise-grade documentation, run as student-leadership infrastructure.

Team Portfolio

Website · 2025A sleek team portfolio site.

A clean, minimalist team-portfolio website built with Next.js 16, Tailwind CSS, and Anime.js — responsive, animated, and fast.

Balance My Life

Game · SUTD · 2023A decision-making life-balance game in Python.

A top-down life-sim built from scratch on a Python/tkinter canvas for SUTD: juggle resources across a daily schedule, move through scenarios with NPCs and a cat companion, and make choices that keep a virtual life in balance — sprites, a HUD, clock, and progress bars all hand-built, running at 60 FPS.

Ainsley Woo
♦ Selected Work
AgentGuard
SentinelLM
origin
AISOCA
QASP
Malware Analysis Sandbox
AI Attack Validation Suite
BenchAudit
HELIX
JANUS
VAC
CALYPSO
RLP
AINSLEY
Beyond the Reflexion Ceiling
Predictive Vulnerability Engine
Zero-Trust Attack-Path Mapper
Autonomous Malware Attribution
Continuous Supply-Chain Trust Verifier
NMAT
Hallucinations at the Firewall
Prompt Refiner
Agendaler
Sports Month 2026
Team Portfolio
Balance My Life
AgentGuard
Agent security
SentinelLM
LLM security
origin
Agentic coding harness
AISOCA
AI SOC analyst
QASP
Post-quantum networking
Malware Analysis Sandbox
Malware analysis
AI Attack Validation Suite
AI red-teaming
BenchAudit
Eval integrity
HELIX
LLM-agent memory
JANUS
Theory-of-mind / belief
VAC
Multi-agent trust
CALYPSO
World models / planning
RLP
Latent planning
AINSLEY
Efficient ML inference
Beyond the Reflexion Ceiling
Agent training
Predictive Vulnerability Engine
Post-quantum crypto
Zero-Trust Attack-Path Mapper
Zero-trust security
Autonomous Malware Attribution
Malware analysis
Continuous Supply-Chain Trust Verifier
Supply-chain security
NMAT
Network analysis
Hallucinations at the Firewall
AI safety · AAAI 2026
Prompt Refiner
Web app · prompt engineering
Agendaler
Proptech · SG real estate
Sports Month 2026
Event platform
Team Portfolio
Website
Balance My Life
Game · SUTD
Scroll