Instructions to use NANI-Nithin/granite-4.2-8b-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use NANI-Nithin/granite-4.2-8b-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf NANI-Nithin/granite-4.2-8b-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf NANI-Nithin/granite-4.2-8b-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf NANI-Nithin/granite-4.2-8b-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf NANI-Nithin/granite-4.2-8b-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf NANI-Nithin/granite-4.2-8b-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf NANI-Nithin/granite-4.2-8b-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf NANI-Nithin/granite-4.2-8b-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf NANI-Nithin/granite-4.2-8b-GGUF:Q4_K_M
Use Docker
docker model run hf.co/NANI-Nithin/granite-4.2-8b-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use NANI-Nithin/granite-4.2-8b-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "NANI-Nithin/granite-4.2-8b-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "NANI-Nithin/granite-4.2-8b-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/NANI-Nithin/granite-4.2-8b-GGUF:Q4_K_M
- Ollama
How to use NANI-Nithin/granite-4.2-8b-GGUF with Ollama:
ollama run hf.co/NANI-Nithin/granite-4.2-8b-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use NANI-Nithin/granite-4.2-8b-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf NANI-Nithin/granite-4.2-8b-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "NANI-Nithin/granite-4.2-8b-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use NANI-Nithin/granite-4.2-8b-GGUF with Docker Model Runner:
docker model run hf.co/NANI-Nithin/granite-4.2-8b-GGUF:Q4_K_M
- Lemonade
How to use NANI-Nithin/granite-4.2-8b-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull NANI-Nithin/granite-4.2-8b-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.granite-4.2-8b-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use NANI-Nithin/granite-4.2-8b-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf NANI-Nithin/granite-4.2-8b-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default NANI-Nithin/granite-4.2-8b-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use NANI-Nithin/granite-4.2-8b-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf NANI-Nithin/granite-4.2-8b-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "NANI-Nithin/granite-4.2-8b-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Granite 4.2 8B · GGUF Quantizations
Ready-to-run GGUF quantizations of IBM Granite 4.2 8B for local inference with llama.cpp, Ollama, LM Studio, Open WebUI, and other GGUF-compatible runtimes.
This repository provides a complete range of quantization formats, from near-lossless BF16 and Q8 models to highly compressed IQ variants for memory-constrained devices.
About the Base Model
Granite 4.2 8B is IBM's dense reasoning language model designed for agentic workflows, coding, tool use, structured outputs, and enterprise-grade reasoning.
Key Properties
| Property | Value |
|---|---|
| Base Model | ibm-granite/granite-4.2-8b |
| Parameters | ~8B |
| Architecture | GraniteForCausalLM |
| Context Window | 128K tokens |
| Reasoning Mode | Switchable thinking / non-thinking |
| Tool Calling | Supported |
| License | Apache 2.0 |
| Format | GGUF |
Features
- 🧠 Native reasoning with optional
<think>mode - 🛠️ Tool calling and agent workflows
- 💻 Strong coding and software engineering capabilities
- 📊 Structured JSON generation
- ⚡ Compatible with llama.cpp and GGUF ecosystems
- 📱 Efficient local deployment on consumer hardware
- 🌍 Multilingual support
Recommended Quantizations
| Quant | Recommendation |
|---|---|
| Q6_K | Highest quality without full precision |
| Q5_K_M | Excellent balance of quality and VRAM |
| Q4_K_M | Default recommendation for most users |
| IQ4_XS | Compact deployment with good quality |
| Q3_K_M | Low-memory environments |
| Q2_K | Maximum compression |
Usage with llama.cpp
./llama-cli \
-m granite-4.2-8b-Q4_K_M.gguf \
-c 8192 \
-p "Explain transformer attention."
Prompt Format
<|im_start|>system
You are a helpful assistant.
<|im_end|>
<|im_start|>user
Explain transformer attention.
<|im_end|>
<|im_start|>assistant
Hardware Guidance
| Quant | Approximate RAM / VRAM |
|---|---|
| Q2_K | 4 GB |
| Q3_K_M | 5 GB |
| Q4_K_M | 6 GB |
| Q5_K_M | 7 GB |
| Q6_K | 9 GB |
| Q8_0 | 11 GB |
| BF16 | 18+ GB |
Actual memory consumption depends on context length, batch size, and runtime configuration.
Intended Use
This model is suitable for:
- Conversational AI
- Coding assistants
- Software engineering agents
- Retrieval-Augmented Generation (RAG)
- Local AI deployments
- Research and experimentation
- Tool-calling workflows
Limitations
- Quantization may slightly reduce model quality compared to the original BF16 checkpoint.
- Smaller quantizations prioritize memory efficiency over accuracy.
- Model outputs may contain inaccuracies or hallucinations.
- Performance depends heavily on the selected quantization level and hardware.
License
This repository distributes quantized versions of IBM Granite 4.2 8B.
The original model is licensed under Apache 2.0. Please refer to the upstream model card for complete licensing and usage information.
Acknowledgements
- IBM Granite Team for the original Granite 4.2 model.
- llama.cpp contributors for GGUF support and quantization tooling.
- Hugging Face for model hosting and distribution.
Citation
If you use these quantizations in research or production, please cite the original Granite 4.2 model from IBM.
- Downloads last month
- 1,039
2-bit
3-bit
4-bit
5-bit
6-bit
16-bit
Model tree for NANI-Nithin/granite-4.2-8b-GGUF
Base model
ibm-granite/granite-4.1-8b-base