다국어 LLM 탈옥 취약점 탐지·방어 연구 Multilingual LLM Jailbreak Vulnerability Detection and Defense
레드팀의 취약점 탐지와 블루팀의 방어를 퍼플팀 체계로 연결해, 윤리적인 다국어 LLM을 위한 문맥 기반 탈옥 가드레일을 구축하는 연구.
A purple-team research program connecting red-team vulnerability discovery with blue-team defenses to build context-aware jailbreak guardrails for ethical multilingual LLMs.
- 연구 목표Research objective
- 윤리적인 다국어 LLM을 위한 문맥 기반 탈옥 가드레일 구축Context-aware jailbreak guardrails for ethical multilingual LLMs
- 주관Funding program
- 한국연구재단 우수신진연구 (과학기술정보통신부)National Research Foundation of Korea, Excellent Young Researchers Program (Ministry of Science and ICT)
- 역할Role
- 참여연구원Participating researcher
- 전체 과제 기간Project period
- 2025.03 – 2028.02
- 참여 기간My participation
- 2025.09 – 2027.08 (종료 예정)(planned end)
- 공식 페이지Official page
- DAMILAB project page ↗
문제와 연구 목표Problem and Research Goals
LLM의 활용이 확대되면서 유해 응답, 보안 악용, 탈옥 공격에 대한 대응이 중요해지고 있습니다. 기존 안전성 연구는 영어 중심의 데이터와 단일 응답 기준의 평가에 의존하는 경우가 많아, 다국어 환경과 맥락(Context)에 따라 달라지는 공격을 충분히 다루기 어렵습니다. 이 과제는 다국어 유해·무해 데이터와 맥락 기반 평가 체계를 마련하고, 화이트박스·블랙박스 공격 연구를 통해 확인한 취약성을 가드레일 개발에 반영하는 것을 목표로 합니다.
Wider LLM adoption raises concerns about harmful responses, misuse, and jailbreak attacks. Safety research that relies primarily on English data and single-response evaluation can miss vulnerabilities that depend on language and context. The project aims to build multilingual harmful and benign datasets, design context-aware evaluation, and use findings from white-box and black-box attack research to inform guardrail development.
과제의 연차별 계획Project Plan by Year
데이터와 평가 체계 구축에서 공격 분석, 가드레일 개발로 이어지는 연차별 연구 계획입니다.
The annual plan progresses from datasets and evaluation to attack analysis and guardrail development.
- 1차년도 — 다국어 데이터셋 및 평가 프레임워크. 안전한 LLM을 위한 다국어 유해·무해 데이터셋을 구축하고, 맥락을 고려한 응답 평가 체계를 개발합니다.
- 2차년도 — 화이트박스·블랙박스 공격 방법론. 모델 내부 정보 접근 여부에 따른 탈옥 공격을 연구하고, 기존 방식이 놓치는 취약성을 분석합니다.
- 3차년도 — 최적화된 가드레일 프레임워크. 데이터·평가·공격 연구 결과를 연결하여 다국어 LLM을 위한 가드레일을 개발합니다.
- Year 1 — Multilingual datasets and evaluation. Build multilingual harmful and benign datasets and develop a framework for evaluating responses in context.
- Year 2 — White-box and black-box attack methods. Study jailbreak methods under different levels of model access and analyze vulnerabilities overlooked by existing approaches.
- Year 3 — Optimized guardrail framework. Connect the dataset, evaluation, and attack findings to develop guardrails for multilingual LLMs.
수행 연구: 프롬프트 위치 취약성 분석Research: Positional Vulnerability in Prompts
기존 GCG 기반 탈옥 공격은 주로 프롬프트 끝(suffix)에 adversarial token을 삽입합니다. SlotGCG 연구는 프롬프트 내부의 다른 삽입 위치에도 취약성이 있다는 점에 주목했습니다. 각 위치의 취약도를 정량화하는 Vulnerable Slot Score(VSS)를 정의하고, 취약한 위치에 공격 토큰을 분산 배치한 뒤 GCG 기반 최적화를 수행하는 방법을 제안했습니다.
Existing GCG-based attacks typically place adversarial tokens at the end of a prompt. SlotGCG investigates vulnerabilities at other insertion positions. The study introduces the Vulnerable Slot Score (VSS) to quantify positional vulnerability, distributes adversarial tokens across vulnerable slots, and then applies GCG-based optimization.
결과: 실험에서는 기존 방법 대비 평균 공격 성공률(ASR) 14% 향상, 방어 기법 적용 환경에서 최대 42% 높은 ASR과 더 적은 최적화 단계에서의 수렴을 확인했습니다. 논문은 ICLR 2026에 채택되었습니다. 다국어 가드레일 프레임워크 개발은 과제의 후속 연구 목표입니다.
Results: Experiments show a 14% average improvement in attack success rate (ASR), up to 42% higher ASR under defenses, and convergence with fewer optimization steps compared with existing methods. The paper was accepted to ICLR 2026. The broader multilingual guardrail framework remains a subsequent project objective.
개인 기여My Contributions
과제에 참여연구원으로 참여하며 LLM 안전성 평가, 공격 분석, 방어 연구를 경험했습니다. SlotGCG에서는 공동저자로서 VSS 기반 평가·분석 및 논문 제출 과정 지원에 기여했습니다. 이 경험을 바탕으로 단일 응답뿐 아니라 도구 사용과 행동 실행까지 포함하는 LLM Agent의 안전성을 후속 연구로 발전시켰고, 그 결과인 Will the User Ever Know?에 Co-1st Author로 참여하여 EMNLP 2026 메인 컨퍼런스 게재 승인을 받았으며, 구두 발표(Oral) 논문으로 선정되었습니다.
As a participating researcher, I gained experience across LLM safety evaluation, attack analysis, and defense research. As a SlotGCG co-author, I contributed to VSS-based evaluation and analysis and supported the paper submission process. This experience led to follow-up research on LLM agent safety, extending the focus from generated responses to tool use and action execution. The resulting paper, Will the User Ever Know?, on which I am a Co-1st Author, was accepted to the EMNLP 2026 Main Conference and selected for an oral presentation.
후속 연구 성과: EMNLP 2026 Oral · ICoAFollow-up Research Outcome: EMNLP 2026 Oral · ICoA
도구 사용 Agent의 간접 프롬프트 주입(Indirect Prompt Injection, IPI)을 다룬 후속 연구는 Will the User Ever Know? Covert Indirect Prompt Injection Attacks on Tool-Using LLM Agents 논문으로 이어졌으며, EMNLP 2026 메인 컨퍼런스에 게재 승인되어 구두 발표(Oral) 논문으로 선정되었습니다. 아래는 이 논문의 핵심 문제와 연구 성과입니다.
The follow-up research on Indirect Prompt Injection (IPI) in tool-using agents resulted in Will the User Ever Know? Covert Indirect Prompt Injection Attacks on Tool-Using LLM Agents, accepted to the EMNLP 2026 Main Conference and selected for an oral presentation. The research question and contributions of this paper are summarized below.
핵심 문제: 외부 데이터의 악의적 지시로 Agent가 원치 않는 도구·API 동작을 수행하더라도, 최종 응답이 정상적인 작업 결과만 보여주면 사용자는 공격을 알아차리기 어렵습니다. 이 논문은 공격의 성공 여부뿐 아니라 성공한 공격이 사용자에게 드러나는지를 함께 평가합니다.
Research question: An agent may execute unwanted tool or API actions after following malicious instructions in external data, yet show only a normal task result in its final response. The paper evaluates not only whether an attack succeeds, but also whether the successful attack is revealed to the user.
- 평가 지표: 공격 성공이 최종 응답에 드러나는지를 구분하는 CSR·OSR 지표를 제안했습니다.
- 공격 방법 및 실험: ICoA 공격을 제안하고, AgentDojo의 네 대상 모델에서 가장 강한 기준선 대비 CSR 3.79–12.01%p 향상을 보고했습니다.
- Evaluation metrics: Introduced CSR and OSR to distinguish covert and overt attack successes based on what the final response reveals.
- Attack method and experiments: Proposed ICoA and reported a 3.79–12.01 percentage-point CSR improvement over the strongest baseline across four target models on AgentDojo.
관련 워크숍 논문: What Did You Do Behind My Back?!에서도 은밀한 간접 프롬프트 주입을 연구했으며, Co-1st Author로 참여한 이 논문은 ICML 2026 FAGEN Workshop에 게재 승인되었습니다.
Related workshop paper: What Did You Do Behind My Back?! also studied covert indirect prompt injection. I am a Co-1st Author of this paper, accepted to the ICML 2026 FAGEN Workshop.
관련 논문Related Publications
Oral
Will the User Ever Know? Covert Indirect Prompt Injection Attacks on Tool-Using LLM Agents
EMNLP 2026 Main Conference (Oral)
What Did You Do Behind My Back?! Covert Indirect Prompt Injection on Tool-Using LLM Agents
ICML 2026 Workshop on Failure Modes of Agentic AI (FAGEN)
SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks
The Fourteenth International Conference on Learning Representations (ICLR), 2026
팀Team
DAMILAB (동국대학교), damilab.github.io
DAMILAB (Dongguk University), damilab.github.io