Crowdsourcing Multiple Choice Science Questions
2017/07/19 by Welbl, Johannes, Liu, Nelson F., Gardner, Matt · 112 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Machine Learning (stat.ML)
paper · doi:10.48550/arxiv.1707.06209
Abstract
We present a novel method for obtaining high-quality, domain-targeted multiple choice questions from crowd workers. Generating these questions can be difficult without trading away originality, relevance or diversity in the answer options. Our method addresses these problems by leveraging a large corpus of domain-specific text and a small set of existing questions. It produces model suggestions for document selection and answer distractor choice which aid the human question generation process. With this method we have assembled SciQ, a dataset of 13.7K multiple choice science exam questions (Dataset available at http://allenai.org/data.html). We demonstrate that the method produces in-domain questions by providing an analysis of this new dataset and by showing that humans cannot distinguish the crowdsourced questions from original questions. When using SciQ as additional training data to existing questions, we observe accuracy improvements on real science exams.
Cited by
- Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary Loss
- Conformal Cascade: Distribution-Free Accuracy Guarantees for Multi-Tier LLM Inference
- CuraWeb: Joint Optimization of Quality, Redundancy, and Diversity for Web-Scale Pretraining Data
- Deep Delta Learning
- AraMix: Recycling, Refiltering, and Deduplicating to Deliver the Largest Arabic Pretraining Corpus
- Bolmo: Byteifying the Next Generation of Language Models
- VersatileFFN: Achieving Parameter Efficiency in LLMs via Adaptive Wide-and-Deep Reuse
- SonicMoE: Accelerating MoE with IO and Tile-aware Optimizations
- Exact Flow Linear Attention: Exact Solution from Continuous-Time Dynamics
- Revisiting the Scaling Properties of Downstream Metrics in Large Language Model Training
- Luxical: High-Speed Lexical-Dense Text Embeddings
- Distance Is All You Need: Radial Dispersion for Uncertainty Estimation in Large Language Models
- Reconstructing KV Caches with Cross-layer Fusion For Enhanced Transformers
- Nexus: Higher-Order Attention Mechanisms in Transformers
- Dripper: Token-Efficient Main HTML Extraction with a Lightweight LM
- ROOT: Robust Orthogonalized Optimizer for Neural Network Training
- E3-Pruner: Towards Efficient, Economical, and Effective Layer Pruning for Large Language Models
- AICC: Parse HTML Finer, Make Models Better -- A 7.3T AI-Ready Corpus Built by a Model-Based HTML Parser
- Generalist Foundation Models Are Not Clinical Enough for Hospital Operations
- VocalBench-zh: Decomposing and Benchmarking the Speech Conversational Abilities in Mandarin Context
- Probabilities Are All You Need: A Probability-Only Approach to Uncertainty Estimation in Large Language Models
- Optimal Attention Temperature Enhances In-Context Learning under Distribution Shift
- Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMs
- Next-Latent Prediction Transformers Learn Compact World Models
- Do Not Step Into the Same River Twice: Learning to Reason from Trial and Error
- SciTrust 2.0: A Comprehensive Framework for Evaluating Trustworthiness of Large Language Models in Scientific Applications
- Gaperon: A Peppered English-French Generative Language Model Suite
- Multi-Agent Debate Strategies: Survey, Taxonomy, and Challenges
- Keyless Attention: Value-Space Routing and Value-Only Caching for Efficient Transformers
- Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
- Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
- Enhancing Linguistic Competence of Language Models through Pre-training with Language Learning Tasks
- Language Model Behavioral Phases are Consistent Across Architecture, Training Data, and Scale
- Beyond Line-Level Filtering for the Pretraining Corpora of LLMs
- ChessQA: Evaluating Large Language Models for Chess Understanding
- MR-Align: Meta-Reasoning Informed Factuality Alignment for Large Reasoning Models
- SeeDNorm: Self-Rescaled Dynamic Normalization
- From Slides to Chatbots: Enhancing Large Language Models with University Course Materials
- Designing and Evaluating Hint Generation Systems for Science Education
- ARC-Encoder: learning compressed text representations for large language models
- Context-level Language Modeling by Learning Predictive Context Embeddings
- Data-Centric Lessons To Improve Speech-Language Pretraining
- Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs
- MARS-M: When Variance Reduction Meets Matrices
- ReXMoE: Reusing Experts with Minimal Overhead in Mixture-of-Experts
- EduAdapt: A Question Answer Benchmark Dataset for Evaluating Grade-Level Adaptability in LLMs
- Midtraining Bridges Pretraining and Posttraining Distributions
- Rewriting History: A Recipe for Interventional Analyses to Study Data Effects on Model Behavior
- ESI: Epistemic Uncertainty Quantification via Semantic-preserving Intervention for Large Language Models
- MatSciBench: Benchmarking the Reasoning Ability of Large Language Models in Materials Science
- Deconstructing Attention: Investigating Design Principles for Effective Language Modeling
- ADVICE: Answer-Dependent Verbalized Confidence Estimation
- Trace Length is a Simple Uncertainty Signal in Reasoning Models
- ELAIPBench: A Benchmark for Expert-Level Artificial Intelligence Paper Understanding
- ProxRouter: Proximity-Weighted LLM Query Routing for Improved Robustness to Outliers
- MeSH: Memory-as-State-Highways for Recursive Transformers
- Reusing Overtrained Language Models Saturates Scaling
- Training Dynamics Impact Post-Training Quantization Robustness
- BLISS: A Lightweight Bilevel Influence Scoring Method for Data Selection in Language Model Pretraining
- Stratum: System-Hardware Co-Design with Tiered Monolithic 3D-Stackable DRAM for Efficient MoE Serving
- Measuring Language Model Hallucinations Through Distributional Correctness
- Benchmarking Foundation Models with Retrieval-Augmented Generation in Olympic-Level Physics Problem Solving
- Composer: A Search Framework for Hybrid Neural Architecture Design
- Microsaccade-Inspired Probing: Positional Encoding Perturbations Reveal LLM Misbehaviours
- CAST: Continuous and Differentiable Semi-Structured Sparsity-Aware Training for Large Language Models
- Understanding the Mixture-of-Experts with Nadaraya-Watson Kernel
- The Flaw of Averages: Quantifying Uniformity of Performance on Benchmarks
- Mechanisms of Matter: Language Inferential Benchmark on Physicochemical Hypothesis in Materials Synthesis
- Conda: Column-Normalized Adam for Training Large Language Models Faster
- Negative Pre-activations Differentiate Syntax
- PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Space
- Tracing the Representation Geometry of Language Models from Pretraining to Post-training
- MoE-PHDS: One MoE checkpoint for flexible runtime sparsity
- IA2: Alignment with ICL Activations Improves Supervised Fine-Tuning
- JGU Mainz's Submission to the WMT25 Shared Task on LLMs with Limited Resources for Slavic Languages: MT and QA
- IIET: Efficient Numerical Transformer via Implicit Iterative Euler Method
- COSPADI: Compressing LLMs via Calibration-Guided Sparse Dictionary Learning
- Harnessing the Potential of Optimizing Data Mixtures via Bayesian Domain Reweighting
- Σ-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems
- The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models
- Experience Scaling: Post-Deployment Evolution For Large Language Models
- Enhancing Scientific Visual Question Answering via Vision-Caption aware Supervised Fine-Tuning
- Synthetic bootstrapped pretraining
- Topic Coverage-based Demonstration Retrieval for In-Context Learning
- Automated MCQA Benchmarking at Scale: Evaluating Reasoning Traces as Retrieval Sources for Domain Adaptation of Small Language Models
- GrACE: A Generative Approach to Better Confidence Elicitation in Large Language Models
- Unbiased Reasoning for Knowledge-Intensive Tasks in Large Language Models via Conditional Front-Door Adjustment
- Does This Look Familiar to You? Knowledge Analysis via Model Internal Representations
- Crown, Frame, Reverse: Layer-Wise Scaling Variants for LLM Pre-Training
- CTCC: A Robust and Stealthy Fingerprinting Framework for Large Language Models via Cross-Turn Contextual Correlation Backdoor
- On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
- Drivel-ology: Challenging LLMs with Interpreting Nonsense with Depth
- EverTracer: Hunting Stolen Large Language Models via Stealthy and Robust Probabilistic Fingerprint
- Implicit Reasoning in Large Language Models: A Comprehensive Survey
- QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting
- Unlocking the Effectiveness of LoRA-FP for Seamless Transfer Implantation of Fingerprints in Downstream Models
- Universal Properties of Activation Sparsity in Modern Large Language Models
- Decoding Memories: An Efficient Pipeline for Self-Consistency Hallucination Detection
- PSO-Merging: Merging Models Based on Particle Swarm Optimization
- Predicting the Order of Upcoming Tokens Improves Language Modeling
- UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning
- Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap
- Exploiting Vocabulary Frequency Imbalance in Language Model Pre-training
- Sycophancy under Pressure: Evaluating and Mitigating Sycophantic Bias via Adversarial Dialogues in Scientific QA
- Maximum Score Routing For Mixture-of-Experts
- Is GPT-OSS Good? A Comprehensive Evaluation of OpenAI's Latest Open Source Models
- Copyright Protection for Large Language Models: A Survey of Methods, Challenges, and Trends
- BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining
- TiMoE: Time-Aware Mixture of Language Experts
- Tackling Distribution Shift in LLM via KILO: Knowledge-Instructed Learning for Continual Adaptation
- FPEdit: Robust LLM Fingerprinting through Localized Parameter Editing
- MolReasoner: Toward Effective and Interpretable Reasoning for Molecular LLMs
Related