Where model behavior forms
Research direction, the lifecycle it instruments, and the published work it sits beside.
- Channels
- 10 active
- Lifecycle stages
- 7 instrumented
- Reading list
- 6 by others
- Own papers
- None yet
Research channels
- CH01
Trustworthy machine learning
Making model behavior predictable, auditable, and safe to depend on.
- CH02
RLHF & preference optimization
How preference data and policy optimization shape model behavior.
- CH03
AI alignment & model safety
Training-time attacks and safety evaluation for aligned models.
- CH04
Learning dynamics & influence analysis
Per-update gradient tracing to attribute behavior to training data.
- CH05
LLM security
Jailbreaks, data poisoning, and securing third-party agent extensions.
- CH06
Computer vision
Detection and pose pipelines (YOLO, ViT-Pose) in research and production.
- CH07
Multimodal learning
Models and pipelines that cross text, image, and video.
- CH08
Retrieval-augmented generation
Grounded assistants with vector search and honest citations.
- CH09
Agentic AI systems
Tool-using agents, orchestration graphs, and safe third-party extensions.
- CH10
Scalable model deployment
GPU serving, containerized inference, and reproducible training infra.
The lifecycle under the microscope
I study the whole lifecycle of an aligned model: how behavior forms during preference-based training, how training-time attacks can corrupt it, how influence analysis can audit it, and what changes once a model acts through tools and third-party extensions. These are open problems across the field — the published work below frames the questions.
Framed by the open literature
These papers are by other researchers, listed as the context this work builds on — never as his own. Every arXiv identifier was checked against the arXiv API.
TrainingOuyang et al.NeurIPS 2022
Training language models to follow instructions with human feedback
The modern SFT → reward model → PPO pipeline this diagram depicts.
DataCarlini et al.IEEE S&P 2024
Poisoning Web-Scale Training Datasets is Practical
Poisoning the data that models train on is practical, not hypothetical.
TrainingRando & TramèrICLR 2024
Universal Jailbreak Backdoors from Poisoned Human Feedback
Poisoned preference data can implant jailbreak backdoors during RLHF.
SafetyHubinger et al.arXiv 2024
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Backdoored behavior can survive standard safety training.
SafetyPruthi et al.NeurIPS 2020
Estimating Training Data Influence by Tracing Gradient Descent
Model behavior can be attributed to training data by tracing gradient updates.
AgentsGreshake et al.AISec @ CCS 2023
Tool-using LLM applications inherit an injection attack surface from untrusted content.
Publications
Papers will be listed here as they become citable. Until then, the channels above and the systems on the systems station are the accurate picture of the work.
A Google Scholar profile will be linked alongside the first listed paper.
Looking for the full record? Open the mission log →