Hi, I'm Mason Park 👋
AI Researcher at Yonsei University focusing on LLM Behavior & Interpretability. I study how and why large language models behave differently across context, interaction, and execution conditions, and how these behavioral patterns can be interpreted, evaluated, and controlled.
- Context-dependent and conditional model behavior
- Failure modes in reasoning, pragmatics, tool use, and multi-step interaction
- Interpreting behavioral differences across models, conditions, and interventions
- Behavioral analysis of tool-using agents under execution and control constraints
- Tool-using LLM agents and multi-step execution
- Stateful and trajectory-level evaluation
- Execution control, stopping behavior, and post-completion actions
- Cost-aware analysis of agent success and failure
- Indirect intent and pragmatic reasoning
- Human–model disagreement and evaluator mismatch
- Context-sensitive and ambiguity-aware interaction
- Harmful language mitigation and detoxification
- Conditional rewriting and intervention
- Robust and controllable model behavior
Findings of ACL 2026 · Co-first Author
A multimodal benchmark for understanding indirect speech acts from visual, conversational, and sociopragmatic context.
- Introduced READI, a vision-based pragmatic QA benchmark
- Evaluated indirect intent understanding across English and Korean
- Showed that strong multimodal models still struggle as indirectness increases
arXiv Preprint · First Author · 2026
A modular framework for span-guided multilingual detoxification.
- Combined intensity-aware span detection with conditioned generation
- Studied toxicity–meaning trade-offs across models and languages
- Evaluated English, Mandarin Chinese, and Korean
When Does Span-Guided Detoxification Help? Human Preferences and Evaluator Diagnostics in a Controlled Comparison
arXiv Preprint · 2026
A controlled study of when span guidance improves detoxification under human evaluation.
- Compared guided and unguided rewriting under controlled conditions
- Analyzed disagreement between automatic evaluators and human preferences
- Examined when stronger intervention helps — and when it does not
Languages:
Python · Java · Kotlin · JavaScript· SQL
Modeling / DL Frameworks:
PyTorch · Transformers · PEFT · TensorFlow
LLM Work:
LLaMA · Qwen · GPT · T5/mT5 · KoBART · XLM-R
Tools / Infra:
Docker · Vessl.ai · RunPod · Kaggle
FastAPI · LangChain · LangGraph
Cursor · VSCode
Empirical analysis of failure, execution behavior, cost, and stopping decisions in multi-step LLM agents.
Evaluation pipelines based on adversarial and safety benchmarks.
Ambiguity detection and clarification strategies for legal QA systems.
