AI SAFETY & TRUSTWORTHY AGENTS
Yunseok Lee 이윤석
동국대학교 컴퓨터AI학과 인공지능전공 석사과정
DAMILAB · Data Analysis & Machine Intelligence Lab
M.S. Student, Department of Computer Science and AI
DAMILAB · Dongguk University
소개About
안녕하세요. 저는 동국대학교 컴퓨터AI학과 인공지능전공 석사과정에 재학 중이며, DAMILAB에서 안전하고 신뢰할 수 있는 AI를 연구하고 있습니다. 현재 AI Safety와 LLM Agent를 중심으로, 언어 모델이 외부 지식과 도구를 사용할 때 발생하는 보안·신뢰성 문제를 살펴봅니다.
제 연구는 탈옥 공격, 간접 프롬프트 주입, 적대적 강건성과 안전성 평가를 아우릅니다. 연구를 관통하는 질문은 하나입니다. 실제 환경에서 작동하는 AI의 행동을 어떻게 검증하고, 사용자가 더 신뢰할 수 있는 시스템으로 만들 수 있을까? 이 관점을 위성 기후자료 초해상화와 같은 실세계 머신러닝 문제에도 확장하고 있습니다.
Hello! I am an M.S. student in AI at Dongguk University and a researcher at DAMILAB, where I study safe and trustworthy AI. My current work focuses on AI Safety and LLM Agents, especially the security and reliability challenges that arise when language models use external knowledge and tools.
My research spans jailbreak attacks, indirect prompt injection, adversarial robustness, and safety evaluation. A single question connects this work: how can we verify the behavior of AI systems operating in real-world environments and make them more trustworthy to their users? I also extend this perspective to real-world machine learning tasks such as satellite climate-data super-resolution.
소식News
- Co-1st Author로 참여한 PLCWorld가 arXiv에 공개되었습니다. 폐루프 플랜트 시뮬레이션에서 LLM이 생성한 PLC 프로그램의 작업 성공과 안전 위반을 평가하는 벤치마크입니다.
- Our paper PLCWorld, on which I am a Co-1st Author, is now on arXiv. It benchmarks task success and safety violations in LLM-generated PLC programs through closed-loop plant simulation.
- Co-1st Author로 참여한 논문 Will the User Ever Know?가 EMNLP 2026 메인 컨퍼런스 구두 발표(Oral) 논문으로 선정되었습니다.
- Our paper Will the User Ever Know?, on which I am a Co-1st Author, was selected for an oral presentation at the EMNLP 2026 Main Conference.
- Co-1st Author로 참여한 논문 Will the User Ever Know?가 EMNLP 2026 메인 컨퍼런스에 게재 승인되고 arXiv에 공개되었습니다.
- Our paper Will the User Ever Know?, on which I am a Co-1st Author, was accepted to the EMNLP 2026 Main Conference and is now on arXiv.
- 서울미래인재재단의 2026 AI서울테크연구지원사업에 선정되어 총 2,000만 원의 연구 지원을 받게 되었습니다.
- I was selected for the 2026 AI Seoul Tech Research Scholarship by the Seoul Future Foundation, with total research funding of KRW 20 million.
- Co-1st Author로 참여한 논문 What Did You Do Behind My Back?!이 ICML 2026 FAGEN Workshop에 게재 승인되었습니다.
- Our paper What Did You Do Behind My Back?!, on which I am a Co-1st Author, was accepted to the ICML 2026 FAGEN Workshop.
- 공저자로 참여한 SlotGCG가 ICLR 2026에 게재 승인되었습니다.
- SlotGCG, on which I am a co-author, was accepted to ICLR 2026.
논문Publications
PLCWorld · 프로젝트 보기View project
PLCWorld: Benchmarking LLM-Generated PLC Programs in Closed-Loop Plant Simulation
arXiv preprint, 2026
100개 과제의 폐루프 플랜트 시뮬레이션에서 LLM이 생성한 PLC 프로그램을 실행하고, 작업 성공과 안전 위반을 별도로 평가합니다.
Evaluating LLM-generated PLC programs through closed-loop plant simulation across 100 tasks, measuring task success and safety violations separately.
[BibTeX]
@misc{kim2026plcworld,
title={{PLCWorld}: Benchmarking {LLM}-Generated {PLC} Programs in Closed-Loop Plant Simulation},
author={Yunji Kim and Yunseok Lee and Hyunwoo Seo and Jaerim Choi and Woojin Lee},
year={2026},
eprint={2610.02982},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2610.02982}
}
ICoA · 프로젝트 보기View project
Will the User Ever Know? Covert Indirect Prompt Injection Attacks on Tool-Using LLM Agents
EMNLP 2026 Main Conference (Oral)
도구 사용 Agent의 공격 성공과 사용자에게 드러나는 정도를 구분하여 은밀한 간접 프롬프트 주입을 평가합니다.
Evaluating covert indirect prompt injection by separating attack success from what a tool-using agent reveals to its user.
[BibTeX]
@misc{lee2026will,
title={Will the User Ever Know? Covert Indirect Prompt Injection Attacks on Tool-Using {LLM} Agents},
author={Yunseok Lee and Yunji Kim and Woojin Lee},
year={2026},
eprint={2608.30362},
archivePrefix={arXiv},
primaryClass={cs.AI},
note={EMNLP 2026 Main Conference (Oral)},
url={https://arxiv.org/abs/2608.30362}
}
FAGEN · 포스터 보기 (PDF)View poster (PDF)
What Did You Do Behind My Back?! Covert Indirect Prompt Injection on Tool-Using LLM Agents
ICML 2026 Workshop on Failure Modes of Agentic AI (FAGEN)
사용자가 알아차리기 어려운 도구 사용 LLM Agent의 간접 프롬프트 주입 공격을 연구합니다.
Studying indirect prompt injection attacks on tool-using LLM agents that can go unnoticed by users.
[BibTeX]
@inproceedings{lee2026what,
title={What Did You Do Behind My Back?! Covert Indirect Prompt Injection on Tool-Using {LLM} Agents},
author={Yunseok Lee and Yunji Kim and Woojin Lee},
booktitle={Workshop on Failure Modes of Agentic AI at ICML 2026},
year={2026},
url={https://openreview.net/forum?id=ozJ54HKqBe}
}
SlotGCG · 논문 보기Read paper
SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks
The Fourteenth International Conference on Learning Representations (ICLR), 2026
프롬프트 내부의 위치별 취약성을 분석하고, 취약한 위치에 공격 토큰을 배치하는 탈옥 방법을 제안합니다.
Analyzing positional vulnerabilities inside prompts to place adversarial tokens at vulnerable slots.
[BibTeX]
@inproceedings{jeong2026slotgcg,
title={Slot{GCG}: Exploiting the Positional Vulnerability in {LLM}s for Jailbreak Attacks},
author={Seungwon Jeong and Jiwoo Jeong and Hyeonjin Kim and Yunseok Lee and Woojin Lee},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=Fn2rSOnpNf}
}
골프 스윙 · 논문 보기 (PDF)Golf swing · Read paper (PDF)
Explainable Graph-Based Golf Swing Analysis Integrating Club and Body Keypoints for Ball Flight Outcome Prediction
Applied Sciences, 16(8), 3813, 2026
클럽과 신체 키포인트를 통합한 그래프 신경망으로 공의 비행 결과를 예측하고, Integrated Gradients로 스윙 단계별 주요 키포인트의 기여도를 분석합니다.
Predicting ball flight outcomes from club and body keypoints with graph neural networks, and interpreting keypoint contributions across swing phases using Integrated Gradients.
[BibTeX]
@article{jung2026explainable,
title={Explainable Graph-Based Golf Swing Analysis Integrating Club and Body Keypoints for Ball Flight Outcome Prediction},
author={Seunghyeon Jung and Minseok Kim and Hyeonjin Kim and Seungwon Jeong and Yunseok Lee and Yunji Kim and Hyunse Lee and Seoyoung Hong and Gyumin Choi and Jaerim Choi and Woojin Lee},
journal={Applied Sciences},
volume={16},
number={8},
pages={3813},
year={2026},
doi={10.3390/app16083813},
url={https://doi.org/10.3390/app16083813}
}
연구 과제 및 프로젝트Research & Projects
연구 과제Research Projects
레드팀·블루팀을 연결하는 퍼플팀 체계에서 다국어 데이터, 탈옥 공격, 문맥 기반 가드레일을 연구. SlotGCG(ICLR 2026)의 VSS 실험·평가에 기여했으며, 도구 사용 LLM Agent의 안전성으로 발전시킨 후속 연구는 Co-1st Author로 참여한 EMNLP 2026 메인 컨퍼런스 Oral 논문으로 이어짐.
Research on multilingual safety data, jailbreak attacks, and context-aware guardrails through a purple-team framework connecting red and blue teams. Contributed to VSS experiments and evaluation for SlotGCG (ICLR 2026); participated as a Co-1st Author in follow-up research on tool-using agent safety, selected for an oral presentation at the EMNLP 2026 Main Conference.
연구 내용과 기여Research & contributions
GK2A·GK2B 위성 자료를 결합해 전천 기후자료를 250m로 초해상화하고, 결측 보간과 지상 관측 기반 보정을 수행. 담당 세부 연구에서 40만 건 이상의 데이터 전처리 및 차이학습 구조 제안에 기여하고 2km → 250m 해상도 향상과 1차년도 RMSE 3K 이하 목표를 달성.
Combines GK2A/GK2B satellite data for 250m all-sky climate-data super-resolution, missing-value interpolation, and ground-observation calibration. Contributed preprocessing of 400,000+ records and the calibration design; achieved 2km → 250m resolution and the year-one RMSE target of ≤3K.
연구 내용과 기여Research & contributions산학·개발 프로젝트Industry & Development Projects
작품 이미지·메타데이터를 벡터화하고, Semantic Search와 LLM 기반 설명 생성을 연결한 AI 도슨트 구현. VR 클라이언트와 AI 서버 간 응답 지연을 고려해 RAG 데이터베이스를 구축한 10개월 산학 프로젝트.
Built an AI docent using artwork image/metadata embeddings, semantic retrieval, and LLM-generated explanations. Developed a RAG database with VR client–AI server latency in mind during a ten-month industry–university project.
프로젝트 자세히Project details
IoT·Wi-Fi·사용자 참여 데이터를 통합하여 실시간 밀집도 분석, 위험 구역 탐지, 모바일 신고 기능을 구현. 실제 캠퍼스 테스트베드 데이터를 활용했으며 2024년 해커톤 ICONICTHON(아코톤) 대상과 상금 100만 원 수상.
Integrated IoT, Wi-Fi, and user-submitted data for crowd-density analysis, risk detection, and mobile incident reporting. Used real campus testbed data and received the ICONICTHON 2024 Hackathon Grand Prize with a KRW 1 million cash prize.
프로젝트 자세히Project details
공공데이터와 GPS를 결합한 실시간 위험 구역 알림 및 이동 지원 서비스. TTS와 온디바이스 음성인식을 연결해 음성으로 조작 가능한 사용자 화면을 구현.
Combined public data and GPS for context-aware risk alerts and mobility assistance. Implemented voice-operated interfaces using text-to-speech and on-device speech recognition.
프로젝트 자세히Project details수상 및 장학Awards & Scholarships
학력Education
경력Experience
연락처Contact
연구 협업이나 논의는 언제든 환영합니다. 이메일로 편하게 연락 주세요.
I am always open to research collaborations and discussions. Feel free to reach out.
이메일Email: yslee0005@dgu.ac.kr
사무실: 서울특별시 중구 필동로1길 30
동국대학교 신공학관 5112호
Room 5112, New Engineering Building, Dongguk University, Seoul 04620, Korea
Office: Room 5112, New Engineering Building,
Dongguk University, 30 Pildong-ro 1-gil, Jung-gu, Seoul 04620, Republic of Korea
서울특별시 중구 필동로1길 30 동국대학교 신공학관 5112호