LLM red-team & jailbreak evaluation harness
Automated adversarial testing for LLMs — prompt injection, multi-turn jailbreaks, and data-exfiltration — scored with an LLM judge and run as a reproducible benchmark on inspect-ai.
LLM red-team & jailbreak evaluation harness
Automated adversarial testing for LLMs — prompt injection, multi-turn jailbreaks, and data-exfiltration — scored with an LLM judge and run as a reproducible benchmark on inspect-ai.
Trust & Safety for a location-social product
A threat model, a calibrated “creepy-vs-safe” classifier, and an LLM red-team for prompt-injection-driven location leakage — security applied to a privacy-sensitive setting.
Transformer from scratch + paper reproductions
A GPT-style transformer built and trained from scratch, plus notes and reproductions of key LLM safety and evaluation papers. Learning in public.
Why I'm pointing eight years of security engineering at frontier AI — and writing the whole way.