GPT Sw3 126M is an open-source language model by AI-Sweden-Models. Features: 126m LLM, VRAM: 0.6GB, License: other, LLM Explorer Score: 0.09, Arc: 22, HellaSwag: 29.6, MMLU: 24.5, GSM8K: 0.1.
research, evaluation of Nordic language capabilities
Primary Use Cases:
Research on LLMs in Nordic languages, Validation of model capabilities
Limitations:
Bias, Safety issues, Generation diversity, Hallucination, Possibility of harmful or inappropriate content
Supported Languages
da (proficient), sv (proficient), no (proficient), en (proficient), is (proficient)
Training Details
Data Sources:
Litteraturbanken, The Pile, Diva, PubMed, ArXiv, CodeParrot, Familjeliv, Flashback, Parlai, Pushshift.io Reddit dataset, English Math dataset from DeepMind, Swedish Math dataset, OPUS, Movie scripts, Natural Instructions, P3, Norwegian Colossal Corpus, Danish Gigaword, Icelandic Gigaword, Common Crawl, LES, Multilingual C4, OSCAR, Open Web Text, Various public Swedish website scrapes, JobTech/Arbetsförmedlingen, Wikipedia
Data Volume:
1.1TB of UTF-8 encoded text containing 660M documents with a total of 320B tokens
Methodology:
Pretrained using a causal language modeling (CLM) objective utilizing the NeMo Megatron GPT implementation.