vix.ing · top · new · best · stats · spec

SMILES-Prompting: A Novel Approach to LLM Jailbreak Attacks in Chemical Synthesis

2024/10/21 by Aidan Wong, Wong, Aidan, He Cao +5 · 1 citation
Computer Science · #Cloud Data Security Solutions #Computation and Language (cs.CL) #Digital and Cyber Forensics #FOS: Computer and information sciences #Information and Cyber Security

paper · pdf · doi:10.48550/arxiv.2410.15641

openalex publication_date 2024/10/21 · openalex created_date 2024/11/06 · openalex updated_date 2026/07/28

Abstract

The increasing integration of large language models (LLMs) across various fields has heightened concerns about their potential to propagate dangerous information. This paper specifically explores the security vulnerabilities of LLMs within the field of chemistry, particularly their capacity to provide instructions for synthesizing hazardous substances. We evaluate the effectiveness of several prompt injection attack methods, including red-teaming, explicit prompting, and implicit prompting. Additionally, we introduce a novel attack technique named SMILES-prompting, which uses the Simplified Molecular-Input Line-Entry System (SMILES) to reference chemical substances. Our findings reveal that SMILES-prompting can effectively bypass current safety mechanisms. These findings highlight the urgent need for enhanced domain-specific safeguards in LLMs to prevent misuse and improve their potential for positive social impact.

Cited by

Related