ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
2024/06/18 by GLM, Team, :, Zeng, Aohan +56 · 249 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2406.12793
Abstract
We introduce ChatGLM, an evolving family of large language models that we have been developing over time. This report primarily focuses on the GLM-4 language series, which includes GLM-4, GLM-4-Air, and GLM-4-9B. They represent our most capable models that are trained with all the insights and lessons gained from the preceding three generations of ChatGLM. To date, the GLM-4 models are pre-trained on ten trillions of tokens mostly in Chinese and English, along with a small set of corpus from 24 languages, and aligned primarily for Chinese and English usage. The high-quality alignment is achieved via a multi-stage post-training process, which involves supervised fine-tuning and learning from human feedback. Evaluations show that GLM-4 1) closely rivals or outperforms GPT-4 in terms of general metrics such as MMLU, GSM8K, MATH, BBH, GPQA, and HumanEval, 2) gets close to GPT-4-Turbo in instruction following as measured by IFEval, 3) matches GPT-4 Turbo (128K) and Claude 3 for long context tasks, and 4) outperforms GPT-4 in Chinese alignments as measured by AlignBench. The GLM-4 All Tools model is further aligned to understand user intent and autonomously decide when and which tool(s) touse -- including web browser, Python interpreter, text-to-image model, and user-defined functions -- to effectively complete complex tasks. In practical applications, it matches and even surpasses GPT-4 All Tools in tasks like accessing online information via web browsing and solving math problems using Python interpreter. Over the course, we have open-sourced a series of models, including ChatGLM-6B (three generations), GLM-4-9B (128K, 1M), GLM-4V-9B, WebGLM, and CodeGeeX, attracting over 10 million downloads on Hugging face in the year 2023 alone. The open models can be accessed through https://github.com/THUDM and https://huggingface.co/THUDM.
Cited by
- SmartSnap: Proactive Evidence Seeking for Self-Verifying Agents
- MoFu: Scale-Aware Modulation and Fourier Fusion for Multi-Subject Video Generation
- EmoTrace: An Emotion Trajectory-Centered Framework for Psychological Support Dialogue Generation
- Express Language Modeling
- LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding
- Traceable LLM Reasoning for Fake-Order Fraud Detection
- Not All LLM Reasoning is Visible in the Chain-of-Thought
- Child-Oriented AIGC Video Risk Reviewing: A Benchmark and Knowledge-Supported Iterative Reasoning Framework
- Bridging the Copyright Gap: Do Large Vision-Language Models Recognize and Respect Copyrighted Content?
- Enhancing Zero-Shot Time Series Forecasting in Off-the-Shelf LLMs via Noise Injection
- CascadeInfer: Length-Aware Scheduling of LLM Serving with Low Latency and Load Balancing
- A Large Language Model Based Method for Complex Logical Reasoning over Knowledge Graphs
- DramaBench: A Six-Dimensional Evaluation Framework for Drama Script Continuation
- AncientBench: Towards Comprehensive Evaluation on Excavated and Transmitted Chinese Corpora
- A Benchmark for Ultra-High-Resolution Remote Sensing MLLMs
- ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body
- CAPTURE: A Benchmark and Evaluation for LVLMs in CAPTCHA Resolving
- TAO-Net: Two-stage Adaptive OOD Classification Network for Fine-grained Encrypted Traffic Classification
- Towards Fine-Grained Recognition with Large Visual Language Models: Benchmark and Optimization Strategies
- VLSA: Vision-Language-Action Models with Plug-and-Play Safety Constraint Layer
- Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
- SoMe: A Realistic Benchmark for LLM-based Social Media Agents
- Dual Refinement Cycle Learning: Unsupervised Text Classification of Mamba and Community Detection on Text Attributed Graph
- SCAIL: Towards Studio-Grade Character Animation via In-Context Learning of 3D-Consistent Pose Representations
- SA-IQA: Redefining Image Quality Assessment for Spatial Aesthetics with Multi-Dimensional Rewards
- InEx: Hallucination Mitigation via Introspection and Cross-Modal Multi-Agent Collaboration
- RULER-Bench: Probing Rule-based Reasoning Abilities of Next-level Video Generation Models for Vision Foundation Intelligence
- AskNearby: An LLM-Based Application for Neighborhood Information Retrieval and Personalized Cognitive-Map Recommendations
- ChartAnchor: Chart Grounding with Structural-Semantic Fidelity
- OralGPT-Omni: A Versatile Dental Multimodal Large Language Model
- SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition
- Evaluation of Large Language Models for Numeric Anomaly Detection in Power Systems
- Progress by Pieces: Test-Time Scaling for Autoregressive Image Generation
- A Reason-then-Describe Instruction Interpreter for Controllable Video Generation
- PhysChoreo: Physics-Controllable Video Generation with Part-Aware Semantic Grounding
- Vision-Language Models for Automated 3D PET/CT Report Generation
- VideoPerceiver: Enhancing Fine-Grained Temporal Perception in Video Multimodal Large Language Models
- ChineseVideoBench: Benchmarking Multi-modal Large Models for Chinese Video Question Answering
- MultiPriv: Benchmarking Individual-Level Privacy Reasoning in Vision-Language Models
- R-AVST: Empowering Video-LLMs with Fine-Grained Spatio-Temporal Reasoning in Complex Audio-Visual Scenarios
- On 10x Better Scalability: KV Stores Scale Up KV Cache
- NeuroPath: Neurobiology-Inspired Path Tracking and Reflection for Semantically Coherent Retrieval
- Cog-RAG: Cognitive-Inspired Dual-Hypergraph with Theme Alignment Retrieval-Augmented Generation
- TIP and Polish: Text-Image-Prototype Guided Multi-Modal Generation via Commonality-Discrepancy Modeling and Refinement
- Evaluating from Benign to Dynamic Adversarial: A Squid Game for Large Language Models
- Information Capacity: Evaluating the Efficiency of Large Language Models via Text Compression
- Benchmarking Multi-Step Legal Reasoning and Analyzing Chain-of-Thought Effects in Large Language Models
- RedOne 2.0: Rethinking Domain-specific LLM Post-Training in Social Networking Services
- SugarTextNet: A Transformer-Based Framework for Detecting Sugar Dating-Related Content on Social Media with Context-Aware Focal Loss
- Overview of CHIP 2025 Shared Task 2: Discharge Medication Recommendation for Metabolic Diseases Based on Chinese Electronic Health Records
- Open-World 3D Scene Graph Generation for Retrieval-Augmented Reasoning
- LiveStar: Live Streaming Assistant for Real-World Online Video Understanding
- GSE: Evaluating Sticker Visual Semantic Similarity via a General Sticker Encoder
- Plan of Knowledge: Retrieval-Augmented Large Language Models for Temporal Knowledge Graph Question Answering
- ChiMDQA: Towards Comprehensive Chinese Document QA with Fine-grained Evaluation
- IF-CRITIC: Towards a Fine-Grained LLM Critic for Instruction-Following Evaluation
- VesSAM: Efficient Multi-Prompting for Segmenting Complex Vessel
- GraphChain: Large Language Models for Large-scale Graph Analysis via Tool Chaining
- Diffuse Thinking: Exploring Diffusion Language Models as Efficient Thought Proposers for Reasoning
- Traceable Drug Recommendation over Medical Knowledge Graphs
- MM-OPERA: Benchmarking Open-ended Association Reasoning for Large Vision-Language Models
- TheraMind: A Strategic and Adaptive Agent for Longitudinal Psychological Counseling
- TS-Reasoner: Aligning Time Series Foundation Models with LLM Reasoning
- Forecasting Trajectory-Level Safety Risks in Black-Box Multi-Turn Interactions
- DIRECT: Direct Decoding for Efficient and Aligned Sequence Labeling with Large Language Models
- RWGBench: Evaluating Scholarly Positioning in Related Work Generation
- ComboBench: Can LLMs Manipulate Physical Devices to Play Virtual Reality Games?
- LuxIT: A Luxembourgish Instruction Tuning Dataset from Monolingual Seed Data
- LongWeave: A Long-Form Generation Benchmark Bridging Real-World Relevance and Verifiability
- PFEA: An LLM-based High-Level Natural Language Planning and Feedback Embodied Agent for Human-Centered AI
- MoPHES:Leveraging on-device LLMs as Agent for Mobile Psychological Health Evaluation and Support
- Code Aesthetics with Agentic Reward Feedback
- Batch Speculative Decoding Done Right
- LooGLE v2: Are LLMs Ready for Real World Long Dependency Challenges?
- Chinese Discharge Drug Recommendation in Metabolic Diseases with Large Language Models
- Human-Agent Collaborative Paper-to-Page Crafting
- VAR: Visual Attention Reasoning via Structured Search and Backtracking
- UWBench: A Comprehensive Vision-Language Benchmark for Underwater Understanding
- Glyph: Scaling Context Windows via Visual-Text Compression
- DETree: DEtecting Human-AI Collaborative Texts via Tree-Structured Hierarchical Representation Learning
- DSEBench: A Test Collection for Explainable Dataset Search with Examples
- Benchmarking Multimodal Large Language Models for Face Recognition
- Hierarchical Frequency Tagging Probe (HFTP): A Unified Approach to Investigate Syntactic Structure Representations in Large Language Models and the Human Brain
- VideoLucy: Deep Memory Backtracking for Long Video Understanding
- ODI-Bench: Can MLLMs Understand Immersive Omnidirectional Environments?
- VeriCite: Towards Reliable Citations in Retrieval-Augmented Generation via Rigorous Verification
- Learning Hanzi Character Through VR-Based Mortise-Tenon
- SocioBench: Modeling Human Behavior in Sociological Surveys with Large Language Models
- LogiNumSynth: Synthesizing Joint Logical-Numerical Reasoning Problems for Language Models
- AccurateRAG: A Framework for Building Accurate Retrieval-Augmented Question-Answering Applications
- Self-Supervised Representation Learning with ID-Content Modality Alignment for Sequential Recommendation
- SASER: Stego attacks on open-source LLMs
- Path Drift in Large Reasoning Models:How First-Person Commitments Override Safety
- RIPRAG: Hack a Black-box Retrieval-Augmented Generation Question-Answering System with Reinforcement Learning
- Abductive Preference Learning
- Inflated Excellence or True Performance? Rethinking Medical Diagnostic Benchmarks with Dynamic Evaluation
- CFVBench: A Comprehensive Video Benchmark for Fine-grained Multimodal Retrieval-Augmented Generation
- A Unified Biomedical Named Entity Recognition Framework with Large Language Models
- Serial-Parallel Dual-Path Architecture for Speaking Style Recognition
- Instance Relation Learning Network with Label Knowledge Propagation for Few-shot Multi-label Intent Detection
- SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training
- MMA-ASIA: A Multilingual and Multimodal Alignment Framework for Culturally-Grounded Evaluation
- Membership Inference Attacks on Tokenizers of Large Language Models
- CDTP: A Large-Scale Chinese Data-Text Pair Dataset for Comprehensive Evaluation of Chinese LLMs
- HyperVLA: Efficient Inference in Vision-Language-Action Models via Hypernetworks
- QuantAgents: Towards Multi-agent Financial System via Simulated Trading
- Evaluating LLMs for Demographic-Targeted Social Bias Detection: A Comprehensive Benchmark Study
- AgentRL: Scaling Agentic Reinforcement Learning with a Multi-Turn, Multi-Task Framework
- Towards Sampling Data Structures for Tensor Products in Turnstile Streams
- A Granular Study of Safety Pretraining under Model Abliteration
- Brain-Language Model Alignment: Insights into the Platonic Hypothesis and Intermediate-Layer Advantage
- Benchmarking Foundation Models with Retrieval-Augmented Generation in Olympic-Level Physics Problem Solving
- Are Large Language Models Chronically Online Surfers? A Dataset for Chinese Internet Meme Explanation
- On Predictability of Reinforcement Learning Dynamics for Large Language Models
- Structuring Reasoning for Complex Rules Beyond Flat Representations
- SDA-PLANNER: State-Dependency Aware Adaptive Planner for Embodied Task Planning
- VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
- Reference-Free Rating of LLM Responses via Latent Information
- HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
- Mechanisms of Matter: Language Inferential Benchmark on Physicochemical Hypothesis in Materials Synthesis
- SVAC: Scaling Is All You Need For Referring Video Object Segmentation
- Mapping Overlaps in Benchmarks through Perplexity in the Wild
- Your Dense Retriever is Secretly an Expeditious Reasoner
- Decoupling Reasoning and Perception: An LLM-LMM Framework for Faithful Visual Reasoning
- PSRT: Accelerating LRM-based Guard Models via Prefilled Safe Reasoning Traces
- GeoRef: Referring Expressions in Geometry via Task Formulation, Synthetic Supervision, and Reinforced MLLM-based Solutions
- Bias in the Picture: Benchmarking VLMs with Social-Cue News Images and LLM-as-Judge Assessment
- ConMem: Contribution-Aware Memory for Long-Horizon Manufacturing Inspection Logs
- Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models
- AVAM: Universal Training-free Adaptive Visual Anchoring Embedded into Multimodal Large Language Model for Multi-image Question Answering
- Speculating LLMs' Chinese Training Data Pollution from Their Tokens
- RS3DBench: A Comprehensive Benchmark for 3D Spatial Perception in Remote Sensing
- AECBench: A Hierarchical Benchmark for Knowledge Evaluation of Large Language Models in the AEC Field
- EngiBench: A Benchmark for Evaluating Large Language Models on Engineering Problem Solving
- MedFact: A Large-scale Chinese Dataset for Evidence-based Medical Fact-checking of LLM Responses
- Understanding Post-Training Structural Changes in Large Language Models
- USB-Rec: An Effective Framework for Improving Conversational Recommendation Capability of Large Language Model
- Redefining Experts: Interpretable Decomposition of Language Models for Toxicity Mitigation
- ORIC: Benchmarking Object Recognition under Contextual Incongruity in Large Vision-Language Models
- UPRPRC: Unified Pipeline for Reproducing Parallel Resources -- Corpus from the United Nations
- Listening, Imagining & Refining: A Heuristic Optimized ASR Correction Framework with LLMs
- TDRM: Smooth Reward Models with Temporal Difference for LLM RL and Inference
- See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
- Chinese Court Simulation with LLM-Based Agent System
- LTA-thinker: Latent Thought-Augmented Training Framework for Large Language Models on Complex Reasoning
- JustEva: A Toolkit to Evaluate LLM Fairness in Legal Knowledge Inference
- MindVL: Towards Efficient and Effective Training of Multimodal Large Language Models on Ascend NPUs
- AesBiasBench: Evaluating Bias and Alignment in Multimodal Language Models for Personalized Image Aesthetic Assessment
- When Safe Unimodal Inputs Collide: Optimizing Reasoning Chains for Cross-Modal Safety in Multimodal Large Language Models
- SparseDoctor: Towards Efficient Chat Doctor with Mixture of Experts Enhanced Large Language Models
- DeepDive: Advancing Deep Search Agents with Knowledge Graphs and Multi-Turn RL
- Large Language Models Meet Legal Artificial Intelligence: A Survey
- LightAgent: Production-level Open-source Agentic AI Framework
- Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis
- Bridging the Gap Between Ideal and Real-world Evaluation: Benchmarking AI-Generated Image Detection in Challenging Scenarios
- When FinTech Meets Privacy: Securing Financial LLMs with Differential Private Fine-Tuning
- WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation
- SemSteDiff: Generative Diffusion Model-based Coverless Semantic Steganography Communication
- Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
- ACE-RL: Adaptive Constraint-Enhanced Reward for Long-form Generation Reinforcement Learning
- CANDY: Benchmarking LLMs' Limitations and Assistive Potential in Chinese Misinformation Fact-Checking
- SelfAug: Mitigating Catastrophic Forgetting in Retrieval-Augmented Generation via Distribution Self-Alignment
- A Probabilistic Inference Scaling Theory for LLM Self-Correction
- FlexMUSE: Multimodal Unification and Semantics Enhancement Framework with Flexible interaction for Creative Writing
- JudgeAgent: Knowledge-wise and Dynamic LLM Evaluation with Agent-as-Interviewer
- DaMoC: Efficiently Selecting the Optimal Large Language Model for Fine-tuning Domain Tasks Based on Data and Model Compression
- DPF-CM: A Data Processing Framework with Privacy-Preserving Vector Databases for Chinese Medical LLMs Training and Deployment
- Can Large Language Models Master Complex Card Games?
- OmniDPO: A Preference Optimization Framework to Address Omni-Modal Hallucination
- Evaluating Recabilities of Foundation Models: A Multi-Domain, Multi-Dataset Benchmark
- SUMMA: A Multimodal Large Language Model for Advertisement Summarization
- Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning
- Social Bias in Multilingual Language Models: A Survey
- Dynamic Collaboration of Multi-Language Models based on Minimal Complete Semantic Units
- SentiMM: A Multimodal Multi-Agent Framework for Sentiment Analysis in Social Media
- RETAIL: Towards Real-world Travel Planning for Large Language Models
- ShizhenGPT: Towards Multimodal LLMs for Traditional Chinese Medicine
- Ask Good Questions for Large Language Models
- From Scores to Skills: A Cognitive Diagnosis Framework for Evaluating Financial Large Language Models
- Unintended Misalignment from Agentic Fine-Tuning: Risks and Mitigation
- ComputerRL: Scaling End-to-End Online Reinforcement Learning for Computer Use Agents
- Prompt-Induced Linguistic Fingerprints for LLM-Generated Fake News Detection
- Wisdom of the Crowd: Reinforcement Learning from Coevolutionary Collective Feedback
- RadarQA: Multi-modal Quality Analysis of Weather Radar Forecasts
- AgentCDM: Enhancing Multi-Agent Collaborative Decision-Making via ACH-Inspired Structured Reasoning
- We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning
- MAC: A Live Benchmark for Multimodal Large Language Models in Scientific Understanding
- OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chatbot via Understanding-Driven Spoken Dialogue
- SHALE: A Scalable Benchmark for Fine-grained Hallucination Evaluation in LVLMs
- InteChar: A Unified Oracle Bone Character List for Ancient Chinese Language Modeling
- Fine-grained Video Dubbing Duration Alignment with Segment Supervised Preference Optimization
- Quick on the Uptake: Eliciting Implicit Intents from Human Demonstrations for Personalized Mobile-Use Agents
- Interpreting Fedspeak with Confidence: A LLM-Based Uncertainty-Aware Framework Guided by Monetary Policy Transmission Paths
- FineBadminton: A Multi-Level Dataset for Fine-Grained Badminton Video Understanding
- FEAT: A Multi-Agent Forensic AI System with Domain-Adapted Large Language Model for Automated Cause-of-Death Analysis
- FormCoach: Lift Smarter, Not Harder
- Improved Personalized Headline Generation via Denoising Fake Interests from Implicit Feedback
- LLMs vs. Chinese Anime Enthusiasts: A Comparative Study on Emotionally Supportive Role-Playing
- Can Large Models Fool the Eye? A New Turing Test for Biological Animation
- LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models
- Pruning Large Language Models by Identifying and Preserving Functional Networks
- Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges
- CTTS: Collective Test-Time Scaling
- Uni-Layout: Integrating Human Feedback in Unified Layout Generation and Evaluation
- TIBSTC-CoT: A Multi-Domain Instruction Dataset for Chain-of-Thought Reasoning in Language Models
- ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation
- Forecasting NCAA Basketball Outcomes with Deep Learning: A Comparative Study of LSTM and Transformer Models
- MIHBench: Benchmarking and Mitigating Multi-Image Hallucinations in Multimodal Large Language Models
- Activation-Guided Local Editing for Jailbreaking Attacks
- Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts
- ITDR: An Instruction Tuning Dataset for Enhancing Large Language Models in Recommendations
- DeltaVLM: Interactive Remote Sensing Image Change Analysis via Instruction-guided Difference Perception
- Exploring the Link Between Bayesian Inference and Embodied Intelligence: Toward Open Physical-World Embodied AI Systems
- The Evolution of Video Anomaly Detection: A Unified Framework from DNN to MLLM
- T2VParser: Adaptive Decomposition Tokens for Partial Alignment in Text to Video Retrieval
- STITCH: Simultaneous Thinking and Talking with Chunked Reasoning for Spoken Language Models
- Reframe Your Life Story: Interactive Narrative Therapist and Innovative Moment Assessment with Large Language Models
- SDD: Self-Degraded Defense against Malicious Fine-tuning
- CrossPL: Evaluating Large Language Models on Cross Programming Language Code Generation
- Flora: Effortless Context Construction to Arbitrary Length and Scale
- Can LLMs Solve ASP Problems? Insights from a Benchmarking Study (Extended Version)
- MMESGBench: Pioneering Multimodal Understanding and Complex Reasoning Benchmark for ESG Tasks
- DBMS-LLM Integration Strategies in Industrial and Business Applications: Current Status and Future Challenges
- Input Reduction Enhanced LLM-based Program Repair
- Decoupling Knowledge and Reasoning in LLMs: An Exploration Using Cognitive Dual-System Theory
- From Neurons to Semantics: Evaluating Cross-Linguistic Alignment Capabilities of Large Language Models via Neurons Alignment
- MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs
- Beyond Isolated Dots: Benchmarking Structured Table Construction as Deep Knowledge Extraction
- The Ever-Evolving Science Exam
- A Survey of Deep Learning for Geometry Problem Solving
- ReAL-AD: Towards Human-Like Reasoning in End-to-End Autonomous Driving
- LRCTI: A Large Language Model-Based Framework for Multi-Step Evidence Retrieval and Reasoning in Cyber Threat Intelligence Credibility Verification
- DCR: Quantifying Data Contamination in LLMs Evaluation
- Open-Source LLMs Collaboration Beats Closed-Source LLMs: A Scalable Multi-Agent System
- DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models
- RedOne: Revealing Domain-specific LLM Post-Training in Social Networking Services
- Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework
- A Survey of Large Language Models in Discipline-specific Research: Challenges, Methods and Opportunities
- Findings of the BEA 2025 Shared Task on Pedagogical Ability Assessment of AI-powered Tutors
- Toward Real-World Chinese Psychological Support Dialogues: CPsDD Dataset and a Co-Evolving Multi-Agent System
- Corvid: Improving Multimodal Large Language Models Towards Chain-of-Thought Reasoning
- ConsNoTrainLoRA: Data-driven Weight Initialization of Low-rank Adapters using Constraints
- InvestAlign: Overcoming Data Scarcity in Aligning Large Language Models with Investor Decision-Making Processes under Herd Behavior
- Affective-ROPTester: Capability and Bias Analysis of LLMs in Predicting Retinopathy of Prematurity
- Flipping Knowledge Distillation: Leveraging Small Models' Expertise to Enhance LLMs in Text Matching
- LOOM-Scope: a comprehensive and efficient LOng-cOntext Model evaluation framework
- LCDS: A Logic-Controlled Discharge Summary Generation System Supporting Source Attribution and Expert Review
- Pre-Trained Policy Discriminators are General Reward Models
- Ready Jurist One: Benchmarking Language Agents for Legal Intelligence in Dynamic Environments
Related