LLM EXPLORER 59,358 MODELS INDEXED

Llama 3.1 Nemotron Nano 4B V1.1 by nvidia

By nvidia · 3940 downloads

Llama 3.1 Nemotron Nano 4B V1.1 is an open-source language model by nvidia. Features: 4b LLM, VRAM: 9GB, Context: 128K, License: other, LLM Explorer Score: 0.22.

  Arxiv:2408.11796   Arxiv:2502.00203   Arxiv:2505.00949 Base model:finetune:nvidia/lla... Base model:nvidia/llama-3.1-mi...   Conversational Dataset:nvidia/llama-nemotron-...   En   Endpoints compatible   Llama   Llama-3   Nvidia   Pytorch   Region:us   Safetensors   Sharded   Tensorflow

Llama 3.1 Nemotron Nano 4B V1.1 Parameters and Internals

LLM NameLlama 3.1 Nemotron Nano 4B V1.1
Repository πŸ€—https://huggingface.co/nvidia/Llama-3.1-Nemotron-Nano-4B-v1.1 
Base Model(s)  nvidia/Llama-3.1-Minitron-4B-Width-Base   nvidia/Llama-3.1-Minitron-4B-Width-Base
Model Size4b
Required VRAM9 GB
Updated2026-08-02
Maintainernvidia
Model Typellama
Model Files  5.0 GB: 1-of-2   4.0 GB: 2-of-2
Supported Languagesen
Model ArchitectureLlamaForCausalLM
Licenseother
Context Length131072
Model Max Length131072
Transformers Version4.47.1
Tokenizer ClassPreTrainedTokenizerFast
Vocabulary Size128256
Torch Data Typebfloat16

Quantized Models of the Llama 3.1 Nemotron Nano 4B V1.1

Model
Likes
Downloads
VRAM
...Nemotron Nano 4B V1.1 Bnb 4bit02183 GB
... Nano 4B V1.1 Unsloth Bnb 4bit1663 GB

Best Alternatives to Llama 3.1 Nemotron Nano 4B V1.1

Best Alternatives
Context / RAM
Downloads
Likes
4Bcpt256K / 8.8 GB50
HoldMy4BKTO256K / 8.8 GB50
Xgen Small 4B Instruct R256K / 17.7 GB544
Xgen Small 4B Base R256K / 17.7 GB303
SJT 4B146K / 7.6 GB70
Nemotron W 4b MagLight 0.1128K / 9.2 GB93
Loxa 4B128K / 16 GB80
Nemotron W 4b Halo 0.1128K / 9.2 GB223
Aura 4B128K / 9 GB2414
Impish LLAMA 4B128K / 9 GB5066
Note: green Score (e.g. "73.2") means that the model is better than nvidia/Llama-3.1-Nemotron-Nano-4B-v1.1.