asrt-bench
Fires poisoned inputs at a real agent and reads its actual tool-call trace for misuse — not the text it replies with. pip install asrt-bench, then /run · /diff to catch safety regressions before they ship. Deterministic, no LLM judge.
I'm Sanskar — a self-taught AI security engineer. I work on prompt-injection detection and agent red-teaming: open datasets and classifiers used 15k+ times on Hugging Face, and asrt-bench, which catches agent-safety regressions from real tool-call traces. Based in Raipur, India · open to remote roles & freelance.
Tools & data for AI security — each one, what it actually does.
Fires poisoned inputs at a real agent and reads its actual tool-call trace for misuse — not the text it replies with. pip install asrt-bench, then /run · /diff to catch safety regressions before they ship. Deterministic, no LLM judge.
A self-questioning skill for Claude Code, Codex and friends — it interrogates a design across the domains that matter, resolves what the codebase can answer, and asks you only the decisions that need your call.
A daily-use set of skills for coding agents — authority, memory, verification, hands-on learning — each with a clear trigger and exit condition, so long agent runs stay understandable and verifiable.
Runs MCP servers and agent skills in a disposable, network-off container, then probes them with hostile input — so you see behavior, not the manifest's claims.
Open prompt-injection data & classifiers on Hugging Face.
Binary and multi-axis datasets (intent, technique, severity) with group-aware, zero-leakage splits, plus a DistilBERT specialist series and DeBERTa detectors. Labels disclosed as model-generated. Used across the community for training and benchmarking guardrails.
The ideas under the tools.
A systems taxonomy for closed-loop autonomous evaluation — the research foundation behind ASRT.
Stopping failures in agentic systems: why AI doesn't know when it's done.
On modality in autonomous LLM engineering.
The hardest problem in autonomous loops: teaching a model when it's done.
Why agentic systems fail to stop themselves — the sunk-cost fallacy, explained.
A practical walkthrough — Python, Gemini API, FinBERT. The project that started it all.
Teaching the hard parts, from scratch.
Let's talk about your model's blind spots.
Open to remote AI-security roles and freelance red-teaming / evaluation work. Based in Raipur, India — happy across time zones.