LLM EXPLORER 56,604 MODELS INDEXED

Qwen3 4B Thinking Full Pretrain Mix Low Tweet 1M En GPT by AmberYifan

By AmberYifan · 11 downloads

Qwen3 4B Thinking Full Pretrain Mix Low Tweet 1M En GPT is an open-source language model by AmberYifan. Features: 4b LLM, VRAM: 8.1GB, Context: 256K, License: apache-2.0, LLM Explorer Score: 0.2.

  Autotrain compatible Base model:finetune:qwen/qwen3... Base model:qwen/qwen3-4b-think...   Conversational   Endpoints compatible   Full   Generated from trainer   Llama-factory   Qwen3   Region:us   Safetensors   Sharded   Tensorflow

Qwen3 4B Thinking Full Pretrain Mix Low Tweet 1M En GPT Parameters and Internals

LLM NameQwen3 4B Thinking Full Pretrain Mix Low Tweet 1M En GPT
Repository πŸ€—https://huggingface.co/AmberYifan/qwen3-4b-thinking-full-pretrain-mix-low-tweet-1m-en-gpt 
Base Model(s)  Qwen3 4B Thinking 2507   Qwen/Qwen3-4B-Thinking-2507
Model Size4b
Required VRAM8.1 GB
Updated2025-11-08
MaintainerAmberYifan
Model Typeqwen3
Model Files  5.0 GB: 1-of-2   3.1 GB: 2-of-2   0.0 GB
Model ArchitectureQwen3ForCausalLM
Licenseapache-2.0
Context Length262144
Model Max Length262144
Transformers Version4.52.4
Tokenizer ClassQwen2Tokenizer
Padding Token<|endoftext|>
Vocabulary Size151936
Torch Data Typebfloat16
Errorsreplace

Best Alternatives to Qwen3 4B Thinking Full Pretrain Mix Low Tweet 1M En GPT

Best Alternatives
Context / RAM
Downloads
Likes
Fable Traces256K / 8.1 GB5530209
FastContext 1.0 4B SFT256K / 8.1 GB5735357
Qwen3 4B Instruct 2507256K / 8.1 GB3109972911
GRPO 4 70256K / 8.1 GB50
Lightning 4B256K / 8.1 GB136
FastContext 1.0 4B RL256K / 8.1 GB455961
Qwen3 4B Thinking 2507256K / 8.1 GB291023606
Qwen3 4B Instruct 2507 FP8256K / 5.2 GB110074478
Typhoon2.5 Qwen3 4B256K / 8 GB4898776
Neuron 4B Instruct256K / 8.1 GB3131
Note: green Score (e.g. "73.2") means that the model is better than AmberYifan/qwen3-4b-thinking-full-pretrain-mix-low-tweet-1m-en-gpt.