LLM EXPLORER 59,420 MODELS INDEXED

Llama 3 8B Fixed Special Embedding by imone

By imone · 13 downloads

Llama 3 8B Fixed Special Embedding is an open-source language model by imone. Features: 8b LLM, VRAM: 16.1GB, Context: 8K, License: other, LLM Explorer Score: 0.12.

  Endpoints compatible   Llama   Region:us   Safetensors   Sharded   Tensorflow

Llama 3 8B Fixed Special Embedding Parameters and Internals

Model Type 
Causal Language Model
Additional Notes 
Original special token weights were zero, causing NaN gradients. Adjustments made to avoid this issue.
Input Output 
Performance Tips:
Re-initialize special token weights to avoid NaN gradients by setting them to the mean of other token weights.
LLM NameLlama 3 8B Fixed Special Embedding
Repository πŸ€—https://huggingface.co/imone/Llama-3-8B-fixed-special-embedding 
Model Size8b
Required VRAM16.1 GB
Updated2026-07-23
Maintainerimone
Model Typellama
Model Files  5.0 GB: 1-of-4   5.0 GB: 2-of-4   4.9 GB: 3-of-4   1.2 GB: 4-of-4
Model ArchitectureLlamaForCausalLM
Licenseother
Context Length8192
Model Max Length8192
Transformers Version4.40.0
Tokenizer ClassPreTrainedTokenizerFast
Vocabulary Size128256
Torch Data Typebfloat16

Best Alternatives to Llama 3 8B Fixed Special Embedding

Best Alternatives
Context / RAM
Downloads
Likes
...otron 8B UltraLong 4M Instruct4192K / 32.1 GB1135125
UltraLong Thinking4192K / 16.1 GB23
...a 3.1 8B UltraLong 4M Instruct4192K / 32.1 GB17624
...a 3.1 8B UltraLong 2M Instruct2096K / 32.1 GB8759
...otron 8B UltraLong 2M Instruct2096K / 32.1 GB12418
Cthulhu 8B V1.41048K / 16.1 GB1010
...raLong 1M Instruct Abliterated1048K / 32.1 GB49
...a 3.1 8B UltraLong 1M Instruct1048K / 32.1 GB138729
...otron 8B UltraLong 1M Instruct1048K / 32.1 GB70259
...xis Bookwriter Llama3.1 8B Sft1048K / 16.1 GB254
Note: green Score (e.g. "73.2") means that the model is better than imone/Llama-3-8B-fixed-special-embedding.