Where model behavior forms
- Research areas
- 19 in 4 groups
- Lifecycle stages
- 7 instrumented
- Reading list
- 6 by others
- Own papers
- None yet
Research areas
19 areas across 4 groups — what forms behavior during training, what keeps it in bounds afterwards, what the model is made of, and what happens once it starts using tools.
Training & learning dynamics
How behavior forms while a model is being trained, and how to trace it back.
04 areas
- 01
RLHF & preference optimization
How preference data and policy optimization shape model behavior.
- 02
Fine-tuning & adaptation
Supervised fine-tuning across staged conditions, and what each stage does to behavior.
- 03
Learning dynamics & influence analysis
Per-update gradient tracing to attribute behavior to training data.
- 04
Tokenization
How text becomes tokens, and what that choice costs everything downstream of it.
Safety, security & governance
Keeping trained behavior predictable, permitted, and auditable once it is deployed.
06 areas
- 05
Trustworthy machine learning
Making model behavior predictable, auditable, and safe to depend on.
- 06
AI alignment & model safety
Training-time attacks and safety evaluation for aligned models.
- 07
LLM security
Jailbreaks, data poisoning, and securing third-party agent extensions.
- 08
Guardrails
Runtime constraints that keep a deployed model inside the behavior it is permitted.
- 09
Safety evaluation & benchmarking
Building the evaluations that decide whether a model is safe enough to ship.
- 10
Privacy & data governance
Treating data-governance, privacy, and confidentiality as requirements, not paperwork.
Models, modality & efficiency
What a model is made of, what it can perceive, and what it costs to run.
05 areas
- 11
Quantization & compression
Making models smaller and cheaper to serve without losing the behavior that mattered.
- 12
Diffusion methods
Diffusion models for generation, and the training dynamics underneath them.
- 13
Vision-language models
Models that read an image and a prompt together and answer about both.
- 14
Multimodal learning
Models and pipelines that cross text, image, and video.
- 15
Computer vision
Detection and pose pipelines (YOLO, ViT-Pose) in research and production.
Agentic & retrieval systems
Models put to work: tools, retrieval, grounding, and serving.
04 areas
- 16
Agentic AI systems
Tool-using agents, orchestration graphs, and safe third-party extensions.
- 17
Retrieval-augmented generation
Grounded assistants with vector search and honest citations.
- 18
Question answering
Answering with a language model: what makes an answer correct, complete, and checkable.
- 19
Scalable model deployment
GPU serving, containerized inference, and reproducible training infra.
The lifecycle, instrumented
I study the whole lifecycle of an aligned model: how behavior forms during preference-based training, how training-time attacks can corrupt it, how influence analysis can audit it, and what changes once a model acts through tools and third-party extensions. These are open problems across the field — the published work below frames the questions.
Framed by the open literature
These papers are by other researchers, listed as the context this work builds on — never as his own. Every arXiv identifier was checked against the arXiv API.
TrainingOuyang et al.NeurIPS 2022
Training language models to follow instructions with human feedback
The modern SFT → reward model → PPO pipeline this diagram depicts.
DataCarlini et al.IEEE S&P 2024
Poisoning Web-Scale Training Datasets is Practical
Poisoning the data that models train on is practical, not hypothetical.
TrainingRando & TramèrICLR 2024
Universal Jailbreak Backdoors from Poisoned Human Feedback
Poisoned preference data can implant jailbreak backdoors during RLHF.
SafetyHubinger et al.arXiv 2024
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Backdoored behavior can survive standard safety training.
SafetyPruthi et al.NeurIPS 2020
Estimating Training Data Influence by Tracing Gradient Descent
Model behavior can be attributed to training data by tracing gradient updates.
AgentsGreshake et al.AISec @ CCS 2023
Tool-using LLM applications inherit an injection attack surface from untrusted content.
Publications
Papers will be listed here as they become citable. Until then, the research areas above and the systems in the work index are the accurate picture of the work.
A Google Scholar profile will be linked alongside the first listed paper.
Looking for the full record? See the experience record →