GPT Sw3 356M is an open-source language model by AI-Sweden-Models. Features: 356m LLM, VRAM: 1.6GB, License: other, LLM Explorer Score: 0.11, Arc: 23.6, HellaSwag: 37.1, MMLU: 25.9, GSM8K: 0.2.
Research, Evaluation of Large Language Models in Nordic languages
Limitations:
Bias and safety limitations, Possible content inaccuracies and irrelevance, Generation diversity issues, Potential for generating offensive, inappropriate content
Considerations:
Includes data diversity concerns and requires feedback mechanism for affected individuals.
Supported Languages
languages_supported (da, sv, no, en, is), proficiency_level (fluent)
Training Details
Data Sources:
Books from Litteraturbanken, The Pile, Articles from Diva, The Pile: PubMed, The Pile: ArXiv, Code from Code Parrot: Github, Pushshift.io Reddit dataset, English Math dataset, Swedish Math dataset, Summarization data, OPUS, Movie scripts, Natural Instructions, P3, The Norwegian Colossal Corpus, Danish Gigaword, Icelandic Gigaword, The Pile: Stack Exchange, Web Common Crawl, MC4, OSCAR, Open Web Text, Miscellaneous public Swedish websites, Familjeliv Articles, Public Swedish Job Ads, Wikipedia
Data Volume:
1.1TB UTF-8 encoded text
Methodology:
Pretrained using a causal language modeling objective
Model Architecture:
NeMo Megatron GPT
Responsible Ai Considerations
Fairness:
The model has limitations regarding bias and safety.
Transparency:
Communication and transparency around usage is encouraged.
Mitigation Strategies:
Controlled pre-release; feedback collection from Nordic NLP ecosystem.