Gpt2023 is an open-source language model by crumb. Features: 137m LLM, VRAM: 0.3GB, License: mit, LLM Explorer Score: 0.25, ELO: 1346, Arc: 21.9, HellaSwag: 31.1, MMLU: 25.1, GSM8K: 0.3.
Gpt2023 Benchmarks
nn.n% — How the model compares to the reference models: Anthropic Sonnet 3.5 ("so35"), GPT-4o ("gpt4o") or GPT-4 ("gpt4").
Gpt2023 Parameters and Internals
Model Type
Use Cases
Areas:
Limitations: Lack of awareness of some recent events due to finetuning on a limited dataset
Supported Languages
Training Details
Data Sources: common crawl sites, ArXiv, GitHub
Data Volume:
Methodology: Finetuning on existing GPT-2 model with learning rate adjustments
Context Length:
Training Time:
Hardware Used:
Model Architecture: Transformer-based architecture, left-to-right causal language model
Input Output
Input Format: Text input, up to 1024 tokens
Accepted Modalities:
Output Format:
Performance Tips: Setting a seed can help achieve reproducible results
LLM Name Gpt2023 Repository π€ https://huggingface.co/crumb/gpt2023 Model Size 137m Required VRAM 0.3 GB Updated 2026-08-05 Maintainer crumb Model Type gpt2 Model Files 0.3 GB 0.3 GB Supported Languages en Model Architecture GPT2LMHeadModel License mit Model Max Length 1024 Transformers Version 4.29.0.dev0 Tokenizer Class GPT2Tokenizer Vocabulary Size 50257 Torch Data Type bfloat16 Activation Function gelu_new
Rank the Gpt2023 capabilities
Have you tried this model? Rate its performance β this feedback helps the ML community find the right model for their needs.
Instruction Following and Task Automation
Factuality and Completeness of Knowledge
Censorship and Alignment
Data Analysis and Insight Generation
Text Generation
Text Summarization and Feature Extraction
Code Generation
Multi-Language Support and Translation
Best Alternatives to Gpt2023
Note: green Score (e.g. "73.2 ") means that the model is better than crumb/gpt2023 .
Expand