Provably Correct by Construction

Verifiable RL environments,
generated at scale

TarantuLabs' proprietary engine generates dozens of new environments an hour. The current focus is cybersecurity, the hardest domain to verify: multi-step exploit chains where up to five vulnerabilities must compose correctly for the flag to fall out. Each one is deterministically solvable and automatically verified, verifiably correct, by the intended path, enforced by gates: the chain composes or it doesn't, with no human judgment and no partial credit. The hard part isn't the exploits; it's generating environments whose correctness is provable at scale. The public corpus is on Hugging Face as TarantuBench v2 (~10k single-vuln labs) and v1 (100-lab suite with chains).

TarantuBench v2
~10k

Corpus Labs (v2)

Acceptance-gated single-vulnerability labs on Hugging Face for training and broader eval — binary flags, technique-aware splits, raw + balanced configs.

5

Up to 5 Vulnerabilities

Multi-step chains (v1 and the interactive catalog) that must fire in sequence — for long-horizon tool-use, not just single exploits.

Total Visibility

Every HTTP request, tool call, and reasoning trace is logged. Full per-step telemetry into exactly how an agent reasons, acts, and fails.

Verifiable RL Environments Are the Next Frontier

If you need them built at scale, provably correct, deterministically verified, and checked to work via the intended path through exploit gates, reach out at [email protected] or on LinkedIn.

Research

Proof that the environments are real signal, not just puzzles: studies that train and evaluate models directly on them, with full per-step telemetry. The environments came first; this is what they're for.

Latest Research

How NOT To Fine-Tune an Offensive AI Model

One week of Qwen 2.5 14B LoRA SFT on TarantuBench, every val attempt regressed or flatlined. Final layered SFT: 2/19 vs base 3/19.

View full results →

Environment Catalog

Interactive slice of the v1 suite. Launch any lab in your browser and attempt the exploit yourself — the same flag a model has to extract is the one you submit. For the full ~10k training corpus, use TarantuBench v2.

Category
Difficulty
AI Solve
0 scenarios available