scientific research on large language models, especially interpretability research
Limitations:
not intended for deployment, generating harmful or offensive text, not suitable for translation or generating text in other languages, not fine-tuned for downstream contexts
Additional Notes
Pythia-160M-deduped was trained on the Pile after global deduplication; current model released after retraining addressing hyperparameter discrepancies; 154 checkpoints provided per model.