MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs
2025/07/20 by Zhengyuan Shi, Zhao, Chenchen, Shi, Zhengyuan +37 · 2 citations
Computer Science · Materials Science · #Advanced Graph Neural Networks #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning in Materials Science #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2507.19525
openalex publication_date 2025/07/20 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
The emergence of multimodal large language models (MLLMs) presents promising opportunities for automation and enhancement in Electronic Design Automation (EDA). However, comprehensively evaluating these models in circuit design remains challenging due to the narrow scope of existing benchmarks. To bridge this gap, we introduce MMCircuitEval, the first multimodal benchmark specifically designed to assess MLLM performance comprehensively across diverse EDA tasks. MMCircuitEval comprises 3614 meticulously curated question-answer (QA) pairs spanning digital and analog circuits across critical EDA stages - ranging from general knowledge and specifications to front-end and back-end design. Derived from textbooks, technical question banks, datasheets, and real-world documentation, each QA pair undergoes rigorous expert review for accuracy and relevance. Our benchmark uniquely categorizes questions by design stage, circuit type, tested abilities (knowledge, comprehension, reasoning, computation), and difficulty level, enabling detailed analysis of model capabilities and limitations. Extensive evaluations reveal significant performance gaps among existing LLMs, particularly in back-end design and complex computations, highlighting the critical need for targeted training datasets and modeling approaches. MMCircuitEval provides a foundational resource for advancing MLLMs in EDA, facilitating their integration into real-world circuit design workflows. Our benchmark is available at https://github.com/cure-lab/MMCircuitEval.
Citations
- DeepCircuitX: A Comprehensive Repository-Level Dataset for RTL Code Understanding, Generation, and PPA Analysis
- DeepGate4: Efficient and Effective Representation Learning for Circuit Design at Scale
- SemiKong: Curating, Training, and Evaluating A Semiconductor Industry-Specific Large Language Model
- OpenLLM-RTL: Open Dataset and Benchmark for LLM-Aided Design RTL Generation
- GPT-4o System Card
- AmpAgent: An LLM-based Multi-Agent System for Multi-stage Amplifier Schematic Design from Literature for Process and Performance Porting
- MiniCPM-V: A GPT-4V Level MLLM on Your Phone
- The Llama 3 Herd of Models
- ChipExpert: The Open-Source Integrated-Circuit-Design-Specific Large Language Model
- Customized Retrieval Augmented Generation and Benchmarking for EDA Tool Documentation QA
- DeepGate3: Towards Scalable Circuit Representation Learning
- Qwen2 Technical Report
- AutoBench: Automatic Testbench Generation and Evaluation Using LLMs for HDL Design
- ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- EDA Corpus: A Large Language Model Dataset for Enhanced Interaction with OpenROAD
- Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language Models
- MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
- Large circuit models: opportunities and challenges
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- Yi: Open Foundation Models by 01.AI
- BioXP-0.5B: Explainable Medical-AI via RL-GRPO
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
- DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
- InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
- AutoChip: Automating HDL Generation Using LLM Feedback
- ChipNeMo: Domain-Adapted LLMs for Chip Design
- InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
- VerilogEval: Evaluating Large Language Models for Verilog Code Generation
- SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
- VeriGen: A Large Language Model for Verilog Code Generation
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- MMBench: Is Your Multi-modal Model an All-around Player?
- Kosmos-2: Grounding Multimodal Large Language Models to the World
- DeepGate2: Functionality-Aware Circuit Representation Learning
- ChipGPT: How far are we from natural language hardware design
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
- GPT-4 Technical Report
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- Training language models to follow instructions with human feedback
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
- DeepGate: Learning Neural Representations of Logic Gates
- Language Models are Few-Shot Learners
- HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
- SQuAD: 100,000+ Questions for Machine Comprehension of Text
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Cited by
Related