The model exhibits strong mathematical abilities and improved coding skills. It has agent abilities and can perform multi-turn conversations. It has a 32k sequence length.
Training Details
Data Volume:
8x more data than v0.1
Methodology:
Knowledge distilled from all experts into a single dense model
Context Length:
32000
Input Output
Input Format:
Guanaco chat template
Release Notes
Version:
v0.2
Date:
April 13
Notes:
This model is not a single trained expert, instead it's a compressed MOE model, turning it into a dense 22B model. It has been trained on 8x more data than v0.1.