TLO: Observing failure before the final answer
A logit-only diagnostic that follows refusal and compliance margins at every decoding step, revealing when and how safety weakens even when final ASR looks similar.
View researchAI Safety ResearchSeoul · KR
I study how language models become safe or unsafe during generation—token by token, before the final answer hides the process.
01
Research portfolio
A logit-only diagnostic that follows refusal and compliance margins at every decoding step, revealing when and how safety weakens even when final ASR looks similar.
View researchAn evaluation of how incremental memory and repeated context framing can make long-running conversations prioritize attacker-imposed context over safety policy.
View researchA graph-grounded retrieval system that connects regulation clauses and relation paths to each answer, improving Recall@5 from 59.6% to 79.8%.
View researchA domain pipeline combining QLoRA SFT, hybrid retrieval, and answer evaluation to improve correctness, groundedness, formatting, and uncertainty handling.
View in portfolio02
Approach
Safety is not only a property of the final answer. It is a process that can be observed, measured, and improved while generation unfolds.
Study logit trajectories, early-token behavior, refusal margins, and action traces before they collapse into one outcome metric.
Design evaluations that explain when, why, and how instruction-following and safety failures emerge.
Apply trustworthy evaluation to post-training, retrieval systems, and agentic decisions where failures may appear early.
03
Working notes
04
Profile
I am an undergraduate student at Chung-Ang University, pursuing a Bachelor of Art and Technology and a Bachelor of Science in Cyber Security as a convergence major.
My work connects LLM safety, trustworthy evaluation, benchmark automation, GraphRAG, and applied machine learning. I am especially interested in extending failure observability from final answers to tool use, memory, planning, and other agentic behavior.