The model might not consistently show improved abilities to follow instructions, and it could respond inappropriately or get stuck in loops., Although this model is aligned to human preferences and has been evaluated for performance, it is not guaranteed that it will refrain from generating harmful content exclusively., Caution is urged against relying on this model for production or adjacent use-cases.
Training Details
Data Sources:
Anthropic's HH-RLHF Dataset
Data Volume:
First 100K examples
Methodology:
Direct Preference Optimization (DPO)
Model Architecture:
OpenLLaMA 3B v2 architecture
Input Output
Input Format:
Modified version of the Alpaca prompt template
Performance Tips:
Utilize the Alpaca prompt template to obtain best responses for instruction-related tasks.