Agentic Design Review System
2025/08/14 by Nag, Sayan, Joseph, K J, Goswami, Koustava +2
#Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multiagent Systems (cs.MA) #Multimedia (cs.MM)
paper · doi:10.48550/arxiv.2508.10745
Abstract
Evaluating graphic designs involves assessing it from multiple facets like alignment, composition, aesthetics and color choices. Evaluating designs in a holistic way involves aggregating feedback from individual expert reviewers. Towards this, we propose an Agentic Design Review System (AgenticDRS), where multiple agents collaboratively analyze a design, orchestrated by a meta-agent. A novel in-context exemplar selection approach based on graph matching and a unique prompt expansion method plays central role towards making each agent design aware. Towards evaluating this framework, we propose DRS-BENCH benchmark. Thorough experimental evaluation against state-of-the-art baselines adapted to the problem setup, backed-up with critical ablation experiments brings out the efficacy of Agentic-DRS in evaluating graphic designs and generating actionable feedback. We hope that this work will attract attention to this pragmatic, yet under-explored research direction.
Citations
- EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception
- MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks
- Aurelia: Test-time Reasoning Distillation in Audio-Visual LLMs
- AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs
- Design-o-meter: Towards Evaluating and Refining Graphic Designs
- SymDPO: Boosting In-Context Learning of Large Multimodal Models with Symbol Demonstration Direct Preference Optimization
- GPT-4o System Card
- Can GPTs Evaluate Graphic Design Based on Design Principles?
- VITA: Towards Open-Source Interactive Omni Multimodal LLM
- LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
- InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
- Meerkat: Audio-Visual Large Language Model for Grounding in Space and Time
- video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
- OpenCOLE: Towards Reproducible Automatic Graphic Design Generation
- VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
- MeLFusion: Synthesizing Music from Image and Language Cues using Diffusion Models
- PosterLLaVa: Constructing a Unified Multi-modal Layout Generator with LLM
- Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models
- LayoutLLM: Layout Instruction Tuning with Large Language Models for Document Understanding
- LayoutFlow: Flow Matching for Layout Generation
- Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset
- In-context learning enables multimodal large language models to classify cancer pathology images
- Demystifying Tacit Knowledge in Graphic Design: Characteristics, Instances, Approaches, and Guidelines
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual Scenarios
- OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
- Large Multimodal Agents: A Survey
- AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
- InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning
- A Survey on Hallucination in Large Vision-Language Models
- GPT-4V(ision) is a Generalist Web Agent, if Grounded
- Gemini vs GPT-4V: A Preliminary Comparison and Combination of Vision-Language Models Through Qualitative Cases
- AppAgent: Multimodal Agents as Smartphone Users
- Jack of All Tasks, Master of Many: Designing General-purpose Coarse-to-Fine Vision-Language Model
- TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
- APoLLo: Unified Adapter and Prompt Learning for Vision Language Models
- X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
- MM-Narrator: Narrating Long-form Videos with Multimodal In-Context Learning
- COLE: A Hierarchical Generation Framework for Multi-Layered and Editable Graphic Design
- PG-Video-LLaVA: Pixel Grounding Large Video-Language Models
- LayoutPrompter: Awaken the Design Ability of Large Language Models
- CogVLM: Visual Expert for Pretrained Language Models
- Loop Copilot: Conducting AI Ensembles for Music Generation and Iterative Editing
- MusicAgent: An AI Agent for Music Understanding and Generation with Large Language Models
- MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
- Octopus: Embodied Vision-Language Programmer from Environmental Feedback
- Negative Object Presence Evaluation (NOPE) to Measure Object Hallucination in Vision-Language Models
- Analyzing and Mitigating Object Hallucination in Large Vision-Language Models
- AnyMAL: An Efficient and Scalable Any-Modality Augmented Language Model
- MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
- NExT-GPT: Any-to-Any Multimodal LLM
- MMHQA-ICL: Multimodal In-context Learning for Hybrid Question Answering over Text, Tables and Images
- ImageBind-LLM: Multi-modality Instruction Tuning
- Detecting and Preventing Hallucinations in Large Vision Language Models
- Towards General Text Embeddings with Multi-stage Contrastive Learning
- LISA: Reasoning Segmentation via Large Language Model
- WavJourney: Compositional Audio Creation with Large Language Models
- Autonomous Tester Agent Benchmark
- BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs
- EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the Backbone
- GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest
- AVSegFormer: Audio-Visual Segmentation with Transformer
- Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
- Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
- Kosmos-2: Grounding Multimodal Large Language Models to the World
- OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
- MIMIC-IT: Multi-Modal In-Context Instruction Tuning
- BeAts: Bengali Speech Acts Recognition using Multimodal Attention Fusion
- Few-shot Fine-tuning vs. In-context Learning: A Fair Comparison and Evaluation
- ChatBridge: Bridging Modalities with Large Language Model as a Language Catalyst
- Evaluating Object Hallucination in Large Vision-Language Models
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
- ImageBind: One Embedding Space To Bind Them All
- mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
- AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head
- Visual Instruction Tuning
- Towards Flexible Multi-modal Document Models
- LayoutDM: Discrete Diffusion Model for Controllable Layout Generation
- DLT: Conditioned layout generation with Joint Discrete-Continuous Diffusion Layout Transformer
- PaLM-E: An Embodied Multimodal Language Model
- In-Context Retrieval-Augmented Language Models
- RT-1: Robotics Transformer for Real-World Control at Scale
- Structured Prompting: Scaling In-Context Learning to 1,000 Examples
- VoLTA: Vision-Language Transformer with Weakly-Supervised Local-Feature Alignment
- MaPLe: Multi-modal Prompt Learning
- Audio-Visual Segmentation
- Egocentric Video-Language Pretraining
- Instruction Induction: From Few Examples to Natural Language Task Descriptions
- Flamingo: a Visual Language Model for Few-Shot Learning
- Training language models to follow instructions with human feedback
- MERLOT Reserve: Neural Script Knowledge through Vision and Language and\n Sound
- High-Resolution Image Synthesis with Latent Diffusion Models
- High-Resolution Image Synthesis with Latent Diffusion Models
- Grounded Language-Image Pre-training
- UniTAB: Unifying Text and Box Outputs for Grounded Vision-Language Modeling
- Noisy Channel Language Model Prompting for Few-Shot Text Classification
- CanvasVAE: Learning to Generate Vector Graphic Documents
- Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
- AudioCLIP: Extending CLIP to Image, Text and Audio
- Learning Transferable Visual Models From Natural Language Supervision
- ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday\n Tasks
- UNITER: UNiversal Image-TExt Representation Learning
- Gromov-Wasserstein Alignment of Word Embedding Spaces
- Iterative Bregman Projections for Regularized Transportation Problems
- Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering
- Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding
- Evaluation and Enhancement of Semantic Grounding in Large Vision-Language Models
Related