The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
2018/03/09 by Jonathan Frankle, Michael Carbin · 15 voices · 259 citations
#cs.LG #cs.AI #cs.NE
paper · pdf
Abstract
Neural network pruning techniques can reduce the parameter counts of trained networks by over 90%, decreasing storage requirements and improving computational performance of inference without compromising accuracy. However, contemporary experience is that the sparse architectures produced by pruning are difficult to train from the start, which would similarly improve training performance. We find that a standard pruning technique naturally uncovers subnetworks whose initializations made them capable of training effectively. Based on these results, we articulate the "lottery ticket hypothesis:" dense, randomly-initialized, feed-forward networks contain subnetworks ("winning tickets") that - when trained in isolation - reach test accuracy comparable to the original network in a similar number of iterations. The winning tickets we find have won the initialization lottery: their connections have initial weights that make training particularly effective. We present an algorithm to identify winning tickets and a series of experiments that support the lottery ticket hypothesis and the importance of these fortuitous initializations. We consistently find winning tickets that are less than 10-20% of the size of several fully-connected and convolutional feed-forward architectures for MNIST and CIFAR10. Above this size, the winning tickets that we find learn faster than the original network and reach higher test accuracy.
Cited by
- Adversarial Prompts for Acceptance Collapse in Speculative Decoding
- Neural Feature Governance: Extending Atom Prevalence
- Flash EQ-Linear: Accelerating Equivariant Linear Layers via Group-wise Discrete Fourier Transform
- Stabilizing Native Low-Rank LLM Pretraining
- Requential Coding: Pushing the Limits of Model Compression with Self-Generated Training Data
- Double-Scoring: Reliable Extraction of Strong Lottery Tickets
- SHUFFLESPARSE: Learned Shuffles for Structured Sparse Networks
- Functional Equivalence and Geometric Diversity in Neural Network Approximations: An Empirical Characterization
- Parameter-Efficient Continual Fine-Tuning: A Survey
- Generalized Fisher-Weighted SVD: Scalable Kronecker-Factored Fisher Approximation for Compressing Large Language Models
- It Takes a MAESTRO To Prune Bad Experts
- NIRVANA: Structured Pruning Reimagined for Large Language Model Compression
- CoG-Guided Weight Correction for Fault-Tolerant Deep Neural Networks
- WingSpan: Concurrency and Dependence for Sparse and Structured Tensor Compilers
- When Bigger is Worse: A Practitioner's Guide to Model Selection Under Data Scarcity
- The Information Shadow: Measuring Structural Limits on What Language Models Can Learn
- Constrained Hebbian Learning Supports Efficient Representational Allocation under Structural Constraints
- Grow-Prune-Freeze Networks: Adaptive & Continual Learning Technique for Olfactory Navigation
- Weight Pruning Amplifies Bias: A Multi-Method Study of Compressed LLMs for Edge AI
- Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights
- Your Language Model Secretly Contains Personality Subnetworks
- Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network Architectures
- The Dead Salmons of AI Interpretability
- Is Smaller Always Faster? Tradeoffs in Compressing Self-Supervised Speech Transformers
- The Universal Weight Subspace Hypothesis
- Who Does What in Deep Learning? Multidimensional Game-Theoretic Attribution of Function of Neural Units
- Deep RL Needs Deep Behavior Analysis: Exploring Implicit Planning by Model-Free Agents in Open-Ended Environments
- Brain-Like Processing Pathways Form in Models With Heterogeneous Experts
- Position: Solve Layerwise Linear Models First to Understand Neural Dynamical Phenomena (Neural Collapse, Emergence, Lazy/Rich Regime, and Grokking)
- Lean Unet: A Compact Model for Image Segmentation
- A Free Lunch in LLM Compression: Revisiting Retraining after Pruning
- High-Dimensional Search, Low-Dimensional Solution: Decoupling Optimization from Representation
- Rethinking Fine-Tuning: Unlocking Hidden Capabilities in Vision-Language Models
- The Quest for Winning Tickets in Low-Rank Adapters
- CosineGate: Semantic Dynamic Routing via Cosine Incompatibility in Residual Networks
- SlimEdge: Lightweight Distributed DNN Deployment on Constrained Hardware
- Distributed Sparse Interventions in Language Models
- Retraction-Free Optimization over the Stiefel Manifold for the LoRA Fine-Tuning
- Bridging Efficiency and Safety: Formal Verification of Neural Networks with Early Exits
- FlashVLM: Text-Guided Visual Token Selection for Large Multimodal Models
- Mixture-of-Experts with Gradient Conflict-Driven Subspace Topology Pruning for Emergent Modularity
- Block-Recurrent Dynamics in Vision Transformers
- Can abstract concepts from LLM improve SLM performance?
- Secret mixtures of experts inside your LLM
- Metanetworks as Regulatory Operators: Learning to Edit for Requirement Compliance
- Surrogate Neural Architecture Codesign Package (SNAC-Pack)
- TPV: Parameter Perturbations Through the Lens of Test Prediction Variance
- LiePrune: Lie Group and Quantum Geometric Dual Representation for One-Shot Structured Pruning of Quantum Neural Networks
- SHARe-KAN: Post-Training Vector Quantization for Cache-Resident KAN Inference
- ZK-APEX: Zero-Knowledge Approximate Personalized Unlearning with Executable Proofs
- Complexity of One-Dimensional ReLU DNNs
- Exploring Test-time Scaling via Prediction Merging on Large-Scale Recommendation
- HOLE: Homological Observation of Latent Embeddings for Neural Network Interpretability
- Winning the Lottery by Preserving Network Training Dynamics with Concrete Ticket Search
- Theoretical Compression Bounds for Wide Multilayer Perceptrons
- Neural expressiveness for beyond importance model compression
- Mitigating Catastrophic Forgetting in Target Language Adaptation of LLMs via Source-Shielded Updates
- Learning interpretable surface elasticity properties from bulk properties
- Understanding and Harnessing Sparsity in Unified Multimodal Models
- Parameter Reduction Improves Vision Transformers: A Comparative Study of Sharing and Width Reduction
- Forgetting by Pruning: Data Deletion in Join Cardinality Estimation
- ModHiFi: Identifying High Fidelity predictive components for Model Modification
- Exploiting the Experts: Unauthorized Compression in MoE-LLMs
- E3-Pruner: Towards Efficient, Economical, and Effective Layer Pruning for Large Language Models
- Layer-wise Weight Selection for Power-Efficient Neural Network Acceleration
- Teacher-Guided One-Shot Pruning via Context-Aware Knowledge Distillation
- Automatic Pruning Discovery for Large Language Models
- Weight Concentration Regularization for Improving Pruning Robustness Under High Sparsity
- Dynamic Black-box Backdoor Attacks on IoT Sensory Data
- Weight-sparse transformers have interpretable circuits
- Efficient Mathematical Reasoning Models via Dynamic Pruning and Knowledge Distillation
- Which Sparse Autoencoder Features Are Real? Model-X Knockoffs for False Discovery Rate Control
- Explore and Establish Synergistic Effects Between Weight Pruning and Coreset Selection in Neural Network Training
- Prediction horizon shapes representations in predictive learning
- Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention
- Pruning as Regularization: Sensitivity-Aware One-Shot Pruning in ASR
- Hardware-Aware YOLO Compression for Low-Power Edge AI on STM32U5 for Weeds Detection in Digital Agriculture
- CAMP-HiVe: Cyclic Pair Merging based Efficient DNN Pruning with Hessian-Vector Approximation for Resource-Constrained Systems
- Models Got Talent: Identifying High Performing Wearable Human Activity Recognition Models Without Training
- SiamMM: A Mixture Model Perspective on Deep Unsupervised Learning
- APP: Accelerated Path Patching with Task-Specific Pruning
- Sharp Minima Can Generalize: A Loss Landscape Perspective On Data
- The Strong Lottery Ticket Hypothesis for Multi-Head Attention Mechanisms
- TwIST: Rigging the Lottery in Transformers with Independent Subnetwork Training
- Random Initialization of Gated Sparse Adapters
- AI Progress Should Be Measured by Capability-Per-Resource, Not Scale Alone: A Framework for Gradient-Guided Resource Allocation in LLMs
- Diluting Restricted Boltzmann Machines
- Grokking in the Ising Model
- Safe Screening Rules for Group SLOPE
- It's Much Easier for Neural Networks to learn Game of Life Dynamics with the Right Activation Function: Polynomial Kolmogorov-Arnold Networks
- Emergent Sparsity in Frozen Random CNN Feature Extractors for Deep Reinforcement Learning
- Use and usability: concepts of representation in philosophy, neuroscience, cognitive science, and computer science
- Transformers Can Learn Rules They've Never Seen: Proof of Computation Beyond Interpolation
- Mapping Networks
- Multi-Resolution Model Fusion for Accelerating the Convolutional Neural Network Training
- Kernelized Sparse Fine-Tuning with Bi-level Parameter Competition for Vision Models
- Adaptive Training of INRs via Pruning and Densification
- Spatio-temporal Multivariate Time Series Forecast with Chosen Variables
- Frustratingly Easy Task-aware Pruning for Large Language Models
- Pruning and Quantization Impact on Graph Neural Networks
- A flexible framework for structural plasticity in GPU-accelerated sparse spiking neural networks
- A Survey on Cache Methods in Diffusion Models: Toward Efficient Multi-Modal Generation
- C-SWAP: Explainability-Aware Structured Pruning for Efficient Neural Networks Compression
- S2AP: Score-space Sharpness Minimization for Adversarial Pruning
- Towards Unsupervised Open-Set Graph Domain Adaptation via Dual Reprogramming
- From Local to Global: Revisiting Structured Pruning Paradigms for Large Language Models
- Neuronal Group Communication for Efficient Neural representation
- Pruning Overparameterized Multi-Task Networks for Degraded Web Image Restoration
- Convergence, design and training of continuous-time dropout as a random batch method
- Structured Sparsity and Weight-adaptive Pruning for Memory and Compute efficient Whisper models
- Compressibility Measures Complexity: Minimum Description Length Meets Singular Learning Theory
- SMEC: Rethinking Matryoshka Representation Learning for Retrieval Embedding Compression
- Pruning Cannot Hurt Robustness: Certified Trade-offs in Reinforcement Learning
- Medical Interpretability and Knowledge Maps of Large Language Models
- FOSSIL: Harnessing Feedback on Suboptimal Samples for Data-Efficient Generalisation with Imitation Learning for Embodied Vision-and-Language Tasks
- Robustness and Regularization in Hierarchical Re-Basin
- SQS: Bayesian DNN Compression through Sparse Quantized Sub-distributions
- Shift-Invariant Attribute Scoring for Kolmogorov-Arnold Networks via Shapley Value
- Do We Really Need Permutations? Impact of Width Expansion on Linear Mode Connectivity
- SliceFine: The Universal Winning-Slice Hypothesis for Pretrained Networks
- Where to Begin: Efficient Pretraining via Subnetwork Selection and Distillation
- JAI-1: A Thai-Centric Large Language Model
- GUIDE: Guided Initialization and Distillation of Embeddings
- Mixture of Neuron Experts
- Expand Neurons, Not Parameters
- Categorical Invariants of Learning Dynamics
- PolyKAN: A Polyhedral Analysis Framework for Provable and Approximately Optimal KAN Compression
- The Unseen Frontier: Pushing the Limits of LLM Sparsity with Surrogate-Free ADMM
- CHORD: Customizing Hybrid-precision On-device Model for Sequential Recommendation with Device-cloud Collaboration
- FlexiQ: Adaptive Mixed-Precision Quantization for Latency/Accuracy Trade-Offs in Deep Neural Networks
- A universal compression theory for lottery ticket hypothesis and neural scaling laws
- CAST: Continuous and Differentiable Semi-Structured Sparsity-Aware Training for Large Language Models
- Enhancing Certifiable Semantic Robustness via Robust Pruning of Deep Neural Networks
- Growing Winning Subnetworks, Not Pruning Them: A Paradigm for Density Discovery in Sparse Neural Networks
- Effective Model Pruning
- Interpret, prune and distill Donut : towards lightweight VLMs for VQA on document
- Query Circuits: Explaining How Language Models Answer User Prompts
- A Second-Order Perspective on Pruning at Initialization and Knowledge Transfer
- Evaluating the Robustness of Chinchilla Compute-Optimal Scaling
- Differentiable Sparsity via D-Gating: Simple and Versatile Structured Penalization
- Replication and Information Extraction in a Minimal Agent-Environment Model
- Toward a Theory of Generalizability in LLM Mechanistic Interpretability Research
- COSPADI: Compressing LLMs via Calibration-Guided Sparse Dictionary Learning
- Budgeted Broadcast: An Activity-Dependent Pruning Rule for Neural Network Efficiency
- MonoCon: A general framework for learning ultra-compact high-fidelity representations using monotonicity constraints
- Binary Autoencoder for Mechanistic Interpretability of Large Language Models
- A Recovery Guarantee for Sparse Neural Networks
- Smaller is Better: Enhancing Transparency in Vehicle AI Systems via Pruning
- Simplifying Neural Networks During Training
- Dense Supervision, Sparse Updates: On the Sparsity and Geometry of On-Policy Distillation
- Learning from Compressed CT: Feature Attention Style Transfer and Structured Factorized Projections for Resource-Efficient Medical Image Analysis
- SNAC-Pack 2.0: Scaled-Out Surrogate Neural Architecture Codesign
- Efficient Reinforcement Learning by Reducing Forgetting with Elephant Activation Functions
- Do Sparse Subnetworks Exhibit Cognitively Aligned Attention? Effects of Pruning on Saliency Map Fidelity, Sparsity, and Concept Coherence
- TASO: Task-Aligned Sparse Optimization for Parameter-Efficient Model Adaptation
- nDNA -- the Semantic Helix of Artificial Cognition
- Analyzing the Effects of Supervised Fine-Tuning on Model Knowledge from Token and Parameter Levels
- A Novel Differential Feature Learning for Effective Hallucination Detection and Classification
- CoUn: Empowering Machine Unlearning via Contrastive Learning
- Balancing Sparse RNNs with Hyperparameterization Benefiting Meta-Learning
- Towards Privacy-Preserving and Heterogeneity-aware Split Federated Learning via Probabilistic Masking
- Deep Learning-Driven Peptide Classification in Biological Nanopores
- Module-Aware Parameter-Efficient Machine Unlearning on Transformers
- Spontaneous Kolmogorov-Arnold Geometry in Shallow MLPs
- NeuroStrike: Neuron-Level Attacks on Aligned LLMs
- Optimal Brain Restoration for Joint Quantization and Sparsification of LLMs
- Investigating the Lottery Ticket Hypothesis for Variational Quantum Circuits
- Explaining How Quantization Disparately Skews a Model
- Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
- Robust Experts: the Effect of Adversarial Training on CNNs with Sparse Mixture-of-Experts Layers
- Randomness with constraints: constructing minimal models for high-dimensional biology
- GEM: A Scale-Aware and Distribution-Sensitive Sparse Fine-Tuning Framework for Effective Downstream Adaptation
- Application of discrete Ricci curvature in pruning randomly wired neural networks: A case study with chest x-ray classification of COVID-19
- Not All Parameters Are Created Equal: Smart Isolation Boosts Fine-Tuning Performance
- Accelerating the drive towards energy-efficient generative AI with quantum computing algorithms
- Dual-Model Weight Selection and Self-Knowledge Distillation for Medical Image Classification
- WISCA: A Lightweight Model Transition Method to Improve LLM Training via Weight Scaling
- An Empirical Study of Knowledge Distillation for Code Understanding Tasks
- FedUP: Efficient Pruning-based Federated Unlearning for Model Poisoning Attacks
- One Shot vs. Iterative: Rethinking Pruning Strategies for Model Compression
- Neuro-inspired Ensemble-to-Ensemble Communication Primitives for Sparse and Efficient ANNs
- DPad: Efficient Diffusion Language Models with Suffix Dropout
- Z-Pruner: Post-Training Pruning of Large Language Models for Efficiency without Retraining
- ONG: One-Shot NMF-based Gradient Masking for Efficient Model Sparsification
- Quantization vs Pruning: Insights from the Strong Lottery Ticket Hypothesis
- Computational Economics in Large Language Models: Exploring Model Behavior and Incentive Design under Resource Constraints
- Synaptic Pruning: A Biological Inspiration for Deep Learning Regularization
- Towards Scalable Lottery Ticket Networks using Genetic Algorithms
- COMponent-Aware Pruning for Accelerated Control Tasks in Latent Space Models
- Sparsity-Driven Plasticity in Multi-Task Reinforcement Learning
- Recurrent Deep Differentiable Logic Gate Networks
- Optimal Brain Connection: Towards Efficient Structural Pruning
- Task complexity shapes internal representations and robustness in neural networks
- FAIR-Pruner: Leveraging Tolerance of Difference for Flexible Automatic Layer-Wise Neural Network Pruning
- Unifying Mixture of Experts and Multi-Head Latent Attention for Efficient Language Models
- On-Device Diffusion Transformer Policy for Efficient Robot Manipulation
- Dimension reduction with structure-aware quantum circuits for hybrid machine learning
- Beyond topography: Topographic regularization improves robustness and reshapes representations in convolutional neural networks
- Forgetting of task-specific knowledge in model merging-based continual learning
- Reinitializing weights vs units for maintaining plasticity in neural networks
- Uncovering Critical Features for Deepfake Detection through the Lottery Ticket Hypothesis
- FGFP: A Fractional Gaussian Filter and Pruning for Deep Neural Networks Compression
- LinDeps: A Fine-tuning Free Post-Pruning Method to Remove Layer-Wise Linear Dependencies with Guaranteed Performance Preservation
- Universal Neurons in GPT-2: Emergence, Persistence, and Functional Impact
- Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study
- The Right to be Forgotten in Pruning: Unveil Machine Unlearning on Sparse Models
- The Benchmark Illusion: Pruned LLMs Can Pass Multiple Choice but Fail to Answer
- Pruning Increases Orderedness in Recurrent Computation
- Large Learning Rates Simultaneously Achieve Robustness to Spurious Correlations and Compressibility
- Reinforcement Learning Fine-Tunes a Sparse Subnetwork in Large Language Models
- Optimization of DNN-based HSI Segmentation FPGA-based SoC for ADS: A Practical Approach
- Understanding Generalization, Robustness, and Interpretability in Low-Capacity Neural Networks
- Online Training and Pruning of Deep Reinforcement Learning Networks
- Auto-Compressing Networks
- The Emergence of Abstract Thought in Large Language Models Beyond Any Language
- On the Stability of the Jacobian Matrix in Deep Neural Networks
- Olica: Efficient Structured Pruning of Large Language Models without Retraining
- Biological Processing Units: Leveraging an Insect Connectome to Pioneer Biofidelic Neural Architectures
- Flows and Diffusions on the Neural Manifold
- Spatial Lifting for Dense Prediction
- Effects of relational graph modularity and depth on the learning performance of neural networks
- QuarterMap: Efficient Post-Training Token Pruning for Visual State Space Models
- Heterogeneous Graph Prompt Learning via Adaptive Weight Pruning
- Compress Any Segment Anything Model (SAM)
- Interpretability-Aware Pruning for Efficient Medical Image Analysis
- UnIT: Scalable Unstructured Inference-Time Pruning for MAC-efficient Neural Inference on MCUs
- Weight-Space Mixture-of-Experts for Implicit Neural Representation Classification
- Towards Robust Surrogate Models: Benchmarking Machine Learning Approaches to Expediting Phase Field Simulations of Brittle Fracture
- Exploring Sparse Adapters for Scalable Merging of Parameter Efficient Experts
- LoSiA: Efficient High-Rank Fine-Tuning via Subnet Localization and Optimization
- LoRA Fine-Tuning Without GPUs: A CPU-Efficient Meta-Generation Framework for LLMs
- Curvature-Weighted Capacity Allocation: A Minimum Description Length Framework for Layer-Adaptive Large Language Model Optimization
- When Does Pruning Benefit Vision Representations?
- LotteryCodec: Searching the Implicit Representation in a Random Network for Low-Complexity Image Compression
- Diversity-Guided MLP Reduction for Efficient Large Vision Transformers
- SIEDD: Shared-Implicit Encoder with Discrete Decoders
- Not All Explanations for Deep Learning Phenomena Are Equally Valuable
- Masked Gated Linear Unit
- Hyperpruning: Efficient Search through Pruned Variants of Recurrent Neural Networks Leveraging Lyapunov Spectrum
- XTransfer: Modality-Agnostic Few-Shot Model Transfer for Human Sensing at the Edge
- PrunePEFT: Iterative Hybrid Pruning for Parameter-Efficient Fine-tuning of LLMs
- Identifying Speaker Information in Feed-Forward Layers of Self-Supervised Speech Transformers
- Loss-Aware Automatic Selection of Structured Pruning Criteria for Deep Neural Network Acceleration
- Cross-regularization: Adaptive Model Complexity through Validation Gradients
- Binsparse: A Specification for Cross-Platform Storage of Sparse Matrices and Tensors
- Sparse-Reg: Improving Sample Complexity in Offline Reinforcement Learning using Sparsity
- Revisiting LoRA through the Lens of Parameter Redundancy: Spectral Encoding Helps
- Mr. Snuffleupagus at SemEval-2025 Task 4: Unlearning Factual Knowledge from LLMs Using Adaptive RMU
- Sharp Generalization Bounds for Foundation Models with Asymmetric Randomized Low-Rank Adapters
- The Butterfly Effect: Neural Network Training Trajectories Are Highly Sensitive to Initial Conditions
- Dynamic Acoustic Model Architecture Optimization in Training for ASR
- MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs
- Gradient-based Fine-Tuning through Pre-trained Model Regularization
- Compression Aware Certified Training
- Machine Unlearning for Robust DNNs: Attribution-Guided Partitioning and Neuron Pruning in Noisy Environments
- Dynamic Sparse Training of Diagonally Sparse Networks
- SAFE: Finding Sparse and Flat Minima to Improve Pruning
- Neural Network Reprogrammability: A Unified Theme on Model Reprogramming, Prompt Tuning, and Prompt Instruction
- Lottery ticket hypothesis [wikipedia]
Discussions
- The Lottery Ticket Hypothesis: Finding Small, Trainable Neural Networks [hn, 137 points, 14 comments]
- The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks (2018) [hn, 121 points, 27 comments]
- And ML agrees! https://arxiv.org/abs/1803.03635 https://proceedings.neurips.cc/paper/2020/hash/322f62469c5e3c7dc3e58f5a4d1ea399-Abstract.html [bsky, 3 points, 0 comments]
- Oh, very cool. Thank you. arxiv.org/abs/1803.03635 [bsky, 2 points, 0 comments]
- There's … not … much to it. Hmm. My model has about 100,000 parameters. The lottery ticket hypothesis suggests that there could be a super-smart 10,000-ish parameter subnetwork inside it. I think I wa [bsky, 2 points, 1 comments]
- My intuition has long been that this is all related to the lottery ticket hypothesis (arxiv.org/abs/1803.03635) but I've never been able to come up with a good way to show that [bsky, 1 points, 0 comments]
- 8/ Efficiency innovations aren’t theirs alone. Remember when @MIT grad student @jefrankle cracked the code on the Lottery Ticket Hypothesis? AI is immature science; you can have a breakthrough and b [bsky, 1 points, 1 comments]
- Para los fanáticos de los avances tecnológicos, redes neuronales, inteligencia artificial y del futuro en general. "El MIT demostró que puedes eliminar el 90% de una red neuronal sin perder precisión" [bsky, 1 points, 1 comments]
- The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks (2019) [hn, 1 points, 0 comments]
- https://arxiv.org/abs/1803.03635 ニューラルネットワークの枝刈りに関する論文です。 訓練済みのネットワークのパラメータ数を90%以上削減できます。 精度を損なうことなく、ストレージ要件を減らし、推論の計算パフォーマンスを向上させます。 [bsky, 0 points, 0 comments]
- The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks (2018) https:// arxiv.org/abs/1803.03635 # arxiv [mastodon, 0 points, 0 comments]
- The Lottery Ticket Hypothesis: finding sparse trainable NNs with 90% less params #HackerNews https://arxiv.org/abs/1803.03635 [bsky, 0 points, 0 comments]
- The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks (2018) https://arxiv.org/abs/1803.03635 https://news.ycombinator.com/item?id=46470513 [bsky, 0 points, 0 comments]
- The Lottery Ticket Hypothesis: finding sparse trainable NNs with 90% less params https://arxiv.org/abs/1803.03635 [bsky, 0 points, 0 comments]
- オリジナルの Lottery Ticket Hypothesis (https://arxiv.org/abs/1803.03635) のステートメントは添付画像なのでちょっとイメージが違うかもです.いまちゃんと証明されてる版 (https://arxiv.org/abs/2002.00585) は training time を無視して存在だけを議論していて,こっちはわりと自明です(「ランダムに [bsky, 0 points, 1 comments]
Related