LLM EXPLORER 59,420 MODELS INDEXED

Microsoft Phi 3 Mini 128K Instruct HQQ 4bit Smashed by PrunaAI

By PrunaAI · 7 downloads

Microsoft Phi 3 Mini 128K Instruct HQQ 4bit Smashed is an open-source language model by PrunaAI. Features: LLM, VRAM: 2.9GB, Context: 128K, Quantized, Instruction-Based, LLM Explorer Score: 0.12.

  4bit   Custom code   Hqq   Instruct   Phi3   Pruna-ai   Quantized   Region:us

Microsoft Phi 3 Mini 128K Instruct HQQ 4bit Smashed Parameters and Internals

Additional Notes 
Results mentioning 'first' are obtained after the first run of the model... 'Sync' metrics are obtained by syncing all GPU processes and stop measurement when all of them are executed. 'Async' metrics are obtained without syncing all GPU processes and stop when the model output can be used by the CPU.
Input Output 
Performance Tips:
Test the efficiency gains directly in your use-cases.
LLM NameMicrosoft Phi 3 Mini 128K Instruct HQQ 4bit Smashed
Repository πŸ€—https://huggingface.co/PrunaAI/microsoft-Phi-3-mini-128k-instruct-HQQ-4bit-smashed 
Base Model(s)  ORIGINAL_REPO_NAME   /ORIGINAL_REPO_NAME
Required VRAM2.9 GB
Updated2025-09-23
MaintainerPrunaAI
Model Typephi3
Instruction-BasedYes
Model Files  2.9 GB
Quantization Type4bit
Model ArchitecturePhi3ForCausalLM
Context Length131072
Model Max Length131072
Transformers Version4.48.2
Tokenizer ClassLlamaTokenizer
Padding Token<|endoftext|>
Vocabulary Size32064
Torch Data Typebfloat16

Best Alternatives to Microsoft Phi 3 Mini 128K Instruct HQQ 4bit Smashed

Best Alternatives
Context / RAM
Downloads
Likes
...m 128K Instruct 8.0bpw H8 EXL2128K / 13.4 GB74
...m 128K Instruct 6.0bpw H6 EXL2128K / 10.7 GB73
...dium 128K Instruct 8 0bpw EXL2128K / 13.4 GB41
...m 128K Instruct 3.0bpw H6 EXL2128K / 5.6 GB50
...m 128K Instruct 5.0bpw H6 EXL2128K / 8.9 GB50
...28K Instruct Ov Fp16 Int4 Asym128K / 2.5 GB90
...128K Instruct HQQ 2bit Smashed128K / 1.4 GB80
NuExtract Bpw6 EXL24K / 3 GB41
...Mini 4K Geminified 3 0bpw EXL24K / 1.6 GB60
...uct Abliterated V3 2 2bpw EXL24K / 4.2 GB90
Note: green Score (e.g. "73.2") means that the model is better than PrunaAI/microsoft-Phi-3-mini-128k-instruct-HQQ-4bit-smashed.