Cerebras GPT 111M is an open-source language model by cerebras. Features: 111m LLM, VRAM: 0.5GB, License: apache-2.0, LLM Explorer Score: 0.13, Arc: 20.2, HellaSwag: 26.7, MMLU: 25.5.
Cerebras GPT 111M Parameters and Internals
Model Type Transformer-based Language Model, Text Generation, Causal LM
Use Cases
Areas: Research, NLP Applications, Ethics and Alignment Research
Applications: Foundation model for NLP research, Reference implementations
Primary Use Cases: Further research into large language models
Limitations: Not suitable for machine translation tasks, Not tuned for human-facing dialog applications
Considerations: Further safety testing and mitigations should be applied before production use.
Additional Notes Trained and evaluated following the approaches described in the Cerebras-GPT paper.
Training Details
Data Sources:
Data Volume:
Methodology: GPT-3 style architecture with full attention and weight streaming technology
Context Length:
Hardware Used: 16 CS-2 wafer scale systems
Model Architecture:
Responsible Ai Considerations
Fairness: The Pile dataset used has been analyzed for various biases and ethical standpoints.
Mitigation Strategies: Mitigations applied are limited to standard Pile dataset pre-processing.
Input Output
Input Format:
Accepted Modalities:
Output Format:
Rank the Cerebras GPT 111M capabilities
Have you tried this model? Rate its performance β this feedback helps the ML community find the right model for their needs.
Instruction Following and Task Automation
Factuality and Completeness of Knowledge
Censorship and Alignment
Data Analysis and Insight Generation
Text Generation
Text Summarization and Feature Extraction
Code Generation
Multi-Language Support and Translation
Expand