Agent Systems
Evaluation, trace replay, action critique, and tooling around execution agents.
I build evaluation and attribution tools for execution agents, and design benchmarks that make unsafe behavior in embodied systems easier to observe, compare, and discuss.
Current Focus
Evaluation, trace replay, action critique, and tooling around execution agents.
Mechanistic skills, attribution workflows, and model behavior diagnosis.
Unsafe action detection and benchmark design for household scenarios.
Research
Critiquing Execution Agents with Self-Evolving Mechanistic Skills
Critical & Act dual-agent framework, self-evolving interpretability skills, trajectory extraction, and replay tooling.
Unsafe action detection for embodied agents in household scenarios
Benchmark design, video data generation, annotation workflow, and paper figure production.
Survey on large reasoning model attacks
Writing and synthesis around resource-consumption threats in reasoning-heavy language models.
Selected Projects
A small neural-network implementation based on NumPy, covering forward and backward propagation.
Attention and residual implementation on a small-scale named entity recognition task.
Small-scale interpretability exploration on fashion captioning and course experiments.
Notes and experiments on maximum likelihood estimation and probability intuition.
Competition
China Undergraduate Mathematical Contest in Modeling, 2025. The solution focused on UAV smoke-screen occlusion duration and used isolated non-zero occlusion peaks as search initialization.
Contact