Zur Hauptnavigation wechseln Zur Suche wechseln Zum Hauptinhalt wechseln

LANGUAGE MODELS SCALE RELIABLY WITH OVER-TRAINING AND ON DOWNSTREAM TASKS

  • Samir Yitzhak Gadre
  • , Georgios Smyrnis
  • , Vaishaal Shankar
  • , Suchin Gururangan
  • , Mitchell Wortsman
  • , Rulin Shao
  • , Jean Mercat
  • , Alex Fang
  • , Jeffrey Li
  • , Sedrick Keh
  • , Rui Xin
  • , Marianna Nezhurina
  • , Igor Vasiljevic
  • , Jenia Jitsev
  • , Luca Soldaini
  • , Alexandros G. Dimakis
  • , Gabriel Ilharco
  • , Pang Wei Koh
  • , Shuran Song
  • , Thomas Kollar
  • Yair Carmon, Achal Dave, Reinhard Heckel, Niklas Muennighoff, Ludwig Schmidt
  • Columbia University
  • Toyota Research Institute
  • University of Texas at Austin
  • Apple Computer
  • University of Washington
  • Forschungszentrum Jülich (FZJ)
  • LAION
  • Allen Institute for AI
  • University of California at Berkeley
  • Bespoke Labs
  • Stanford University
  • Tel Aviv University
  • Contextual AI

Publikation: Beitrag in Buch/Bericht/KonferenzbandKonferenzbeitragBegutachtung

5 Zitate (Scopus)

Abstract

Scaling laws are useful guides for derisking expensive training runs, as they predict performance of large models using cheaper, small-scale experiments. However, there remain gaps between current scaling studies and how language models are ultimately trained and evaluated. For instance, scaling is usually studied in the compute-optimal training regime (i.e., “Chinchilla optimal” regime). In contrast, models are often over-trained to reduce inference costs. Moreover, scaling laws mostly predict loss on next-token prediction, but models are usually compared on downstream task performance. To address both shortcomings, we create a testbed of 104 models with 0.011B to 6.9B parameters trained with various numbers of tokens on three data distributions. First, we fit scaling laws that extrapolate in both the amount of over-training and the number of model parameters. This enables us to predict the validation loss of a 1.4B parameter, 900B token run (i.e., 32→ over-trained) and a 6.9B parameter, 138B token run (i.e., a compute-optimal run)-each from experiments that take 300→ less compute. Second, we relate the perplexity of a language model to its downstream task performance by proposing a power law. We use this law to predict top-1 error averaged over downstream tasks for the two aforementioned models, using experiments that take 20→ less compute. To facilitate further research on reliable scaling, we provide all results of our experiments. Our experiments are available at https://github.com/mlfoundations/scaling.

OriginalspracheEnglisch
Titel13th International Conference on Learning Representations, ICLR 2025
Herausgeber (Verlag)International Conference on Learning Representations, ICLR
Seiten47297-47318
Seitenumfang22
ISBN (elektronisch)9798331320850
PublikationsstatusVeröffentlicht - 2025
Veranstaltung13th International Conference on Learning Representations, ICLR 2025 - Singapore, Singapur
Dauer: 24 Apr. 202528 Apr. 2025

Publikationsreihe

Name13th International Conference on Learning Representations, ICLR 2025

Konferenz

Konferenz13th International Conference on Learning Representations, ICLR 2025
Land/GebietSingapur
OrtSingapore
Zeitraum24/04/2528/04/25

Fingerprint

Untersuchen Sie die Forschungsthemen von „LANGUAGE MODELS SCALE RELIABLY WITH OVER-TRAINING AND ON DOWNSTREAM TASKS“. Zusammen bilden sie einen einzigartigen Fingerprint.

Dieses zitieren