AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Hugging Face’s transformers library now supports running GGUF quantized checkpoints directly through from_pretrained, reusing llama.cpp’s ggml kernels for performance. Initial support targets Apple Silicon and the Qwen3.5 architecture, with benchmarks benchmarked against llama.cpp itself.

Hugging Face has added native support for GGUF quantized models to its transformers library, letting users run llama.cpp-style quantized checkpoints through the familiar from_pretrained API on their own machines. The feature reuses llama.cpp’s underlying ggml kernels to keep performance close to llama.cpp itself, with initial support focused on Apple Silicon and the Qwen3.5 architecture. The capability is currently available on the transformers main branch, ahead of the next stable release.

According to Hugging Face’s announcement, users can now pick any GGUF checkpoint from the Hub, load it by passing a gguf_file argument to from_pretrained, and generate text with no extra configuration. When weights stay packed on Metal, transformers automatically loads compatible ggml/Metal layer kernels and uses ggml-org/ggml-attn as the attention implementation. If that kernel cannot be fetched, the model falls back to the standard “sdpa” attention with a warning, and users can force sdpa explicitly via attn_implementation="sdpa".

The requirements are specific: an Apple Silicon Mac, a PyTorch version supported by the published ggml-quantization kernel builds (usually the two latest releases), and the latest version of transformers plus a compatible version of the kernels library. Without a compatible quantization kernel, the loader falls back to dequantizing the model, which uses more memory.

Beyond direct model loading, the same GGUF checkpoints can be served through transformers serve, which exposes an OpenAI-compatible API on localhost. Clients such as Jan or Pi can connect by adding a custom OpenAI-compatible provider pointing at that endpoint. Hugging Face states that its reference for local inference performance is llama.cpp, and its benchmark comparison covers three GGUF checkpoints: a small dense model, a larger dense model, and a mixture-of-experts model.

At a glance
announcementWhen: announced April 2026; available via tra…
The developmentHugging Face announced that the transformers library can now run llama.cpp-style GGUF quantized models natively, using ggml kernels for near-llama.cpp performance on Apple Silicon.
Transformers Now Runs Llama.cpp Quants
Local Inference / Hugging Face / April 2026

Transformers Now Runs Llama.cpp Quants

Hugging Face’s transformers library now loads GGUF quantized checkpoints directly through from_pretrained, reusing llama.cpp’s ggml kernels for near-native performance. Initial support targets Apple Silicon and the Qwen3.5 architecture, benchmarked against llama.cpp itself.

8.42 → 2.74 GB
Qwen3.5-4B: BF16 vs Q4_K_M
gguf_file
One argument — zero extra config
3 checkpoints
Benchmarked vs llama.cpp
Apple Silicon
Required hardware (today)
Qwen3.5
First architecture
main branch
Availability — pre-release
ggml-attn
Attention kernel (Metal)
The Development

Two Ecosystems, One API

GGUF — developed by the llama.cpp team — powers tools like Ollama, LM Studio, and Jan, with millions of downloads. Until now, running those checkpoints meant leaving the PyTorch-based transformers stack. This move collapses that boundary: developers keep their code and gain access to quantized checkpoints from Unsloth, LM Studio Community, bartowski, and ggml-org.

01 — Familiar API

No Code Changes

Pick any GGUF checkpoint from the Hub and pass gguf_file to from_pretrained. Text generation works with no extra configuration.

02 — Native Kernels

ggml Performance

When weights stay packed on Metal, transformers loads compatible ggml/Metal layer kernels and uses ggml-org/ggml-attn for attention.

03 — Serving

OpenAI-Compatible API

The same checkpoints run through transformers serve, exposing a localhost endpoint that clients like Jan or Pi connect to as a custom provider.

How It Fits Together

From Checkpoint to Generation

The loading path shows how transformers reuses llama.cpp’s proven machinery — and where it degrades gracefully when pieces are missing.

1

Pick a GGUF file

Any GGUF checkpoint from the Hub — weights, tokenizer, and chat template in a single file.

2

from_pretrained

Pass the gguf_file argument. No extra configuration required.

3

ggml/Metal kernels

Weights stay packed on Metal; ggml-attn handles attention.

4

Fallback: sdpa

If the kernel can’t be fetched, attention falls back to “sdpa” with a warning — force it via attn_implementation=”sdpa”.

Quantization Tradeoffs

Trading Precision for Memory

File sizes for Unsloth’s Qwen3.5-4B show how mixed-precision variants like Q4_K_M — mostly 4-bit weights, sensitive tensors kept higher — shrink models to laptop scale. HF recommends starting at Q4_K_M and stepping up if memory allows.

BF16
8.42 GB
Q6_K
3.53 GB
Q5_K_M
3.14 GB
Q4_K_M
2.74 GB

Quality cost from aggressive quantization “depends on the model and the task” — Hugging Face advises evaluating on your actual workload.

Boundaries

Limits of the Current Rollout

Capability Status Today Timeline
Apple Silicon support✓ AvailableShipped
Qwen3.5 architecture✓ SupportedShipped
CUDA / Linux / Windows✗ Not supportedNo stated timeline
Stable transformers release~ Main branch onlyDate not announced
Other architectures~ UnclearMoE models used in benchmarks
In Their Words

From the Announcement

“We’re adding support for running GGUF models efficiently in transformers, so you can use checkpoints sized for your laptop’s memory through the familiar transformers APIs.”

Hugging Face Announcement

“This is where we are right now. And I’m not gonna lie — it feels pretty magical.”

Julien Chaumond, Co-founder — April 24, 2026
Roadmap

Toward Broader Device Support

Expansion implied along two axes: additional hardware backends and additional model architectures. Users can track the Hub’s GGUF documentation for new quantization types as PyTorch versions update.

💻 Main branch (today)
→
📦 Stable release
→
🎮 CUDA GPUs
→
🧩 More architectures
→
⚡ MoE models
Where I Land

Pragmatic, Well-Executed — But Specific

Rather than reinventing a quantized inference engine, Hugging Face wrapped llama.cpp’s proven ggml kernels in the API its users already know. The value matters most to developers who live in transformers; Ollama and LM Studio already serve everyone else.

What would change the assessment: CUDA support, coverage of architectures beyond Qwen3.5, and independent benchmarks showing generation speed genuinely comparable to llama.cpp on the same hardware. If those arrive within a few release cycles, this becomes a meaningful consolidation of the local AI tooling landscape rather than a convenience feature.

What This Means for Local AI Users

This development matters because it collapses two previously separate ecosystems. GGUF, developed by the llama.cpp team, has become the dominant format for local inference — it powers tools like Ollama, LM Studio, and Jan, and GGUF models have been downloaded millions of times. Until now, running those checkpoints generally meant using llama.cpp-derived tools rather than the PyTorch-based transformers stack.

For developers already building on transformers, this means access to the full range of quantized checkpoints published by Unsloth, LM Studio Community, bartowski, and ggml-org without changing their code. For users with limited hardware, quantization lets a model like Qwen3.5-4B shrink from 8.42 GB in BF16 to 2.74 GB in Q4_K_M, making laptop-scale inference practical. Hugging Face frames the move as making local AI “much easier” for everyday use, a claim echoed by growing interest in local coding agents running mid-sized models on consumer Macs.

Amazon

Top picks for "transformer runs llama"

As an affiliate, we earn on qualifying purchases.

How GGUF Quantization Fits Together

GGUF packages model weights and metadata — including tokenizer information and an optional chat template — in a single file. It supports multiple quantization levels, letting users trade precision for memory footprint. Variants such as Q4_K_M use mixed tensor precision: mostly 4-bit weights while keeping sensitive tensors at higher precision.

Hugging Face’s published file sizes for Unsloth’s Qwen3.5-4B illustrate the tradeoffs: BF16 at 8.42 GB as the unquantized reference, Q6_K at 3.53 GB, Q5_K_M at 3.14 GB, and Q4_K_M at 2.74 GB. The company recommends starting with Q4_K_M and moving to Q5_K_M or Q6_K if more memory is available, while cautioning that the quality cost of aggressive quantization depends on the model and task — users should evaluate on their actual workload.

The timing follows a period of rapid improvement in local inference. Hugging Face co-founder Julien Chaumond recently posted a demonstration of Qwen3.6 27B running inside the Pi coding agent via llama.cpp on a MacBook Pro, writing that for non-trivial tasks on Hugging Face codebases it felt “very, very close” to hitting the latest Claude Opus.

“We’re adding support for running GGUF models efficiently in transformers, so you can use checkpoints sized for your laptop’s memory through the familiar transformers APIs.”

— Hugging Face announcement

Limits of the Current Rollout

Several boundaries remain. Support currently targets Apple Silicon only — there is no stated timeline for CUDA, Linux, or Windows support. Architecture coverage starts with Qwen3.5; it is unclear which additional model families will be added or when. The feature is also only on the transformers main branch for now, and Hugging Face has not announced a date for the next stable release that would include it.

Full benchmark numbers comparing transformers-GGUF performance against llama.cpp across the three test checkpoints were referenced in the announcement but details depend on the specific hardware and models used. Hugging Face itself notes that quality loss from more aggressive quantization “depends on the model and the task” and advises users to evaluate on their own workloads.

Roadmap for Broader Device Support

The immediate next step is the feature shipping in a stable transformers release, removing the need to install from GitHub. Beyond that, Hugging Face’s stated initial focus on Apple Silicon and Qwen3.5 implies likely expansion along two axes: additional hardware backends (such as CUDA GPUs) and additional model architectures, including the mixture-of-experts models already used in its benchmarking. Users can track the Hub’s GGUF documentation for newly supported quantization types and the kernels library for expanded kernel builds as PyTorch versions update.

Where I land

In my view, this is a pragmatic and well-executed move by Hugging Face: rather than reinventing a quantized inference engine, the company wrapped llama.cpp’s proven ggml kernels in the API that its existing user base already knows. The value is real but specific — it matters most to developers who live in transformers and previously had to leave that ecosystem to run laptop-sized quantized models.

The strongest counterargument is that llama.cpp-based tools like Ollama and LM Studio already make local GGUF inference easy for most users, and transformers adds a heavyweight PyTorch dependency that pure llama.cpp installs avoid. For someone not already invested in the transformers stack, this feature may change little. There is also a real risk of it becoming a maintenance burden across PyTorch versions and architectures, given its current single-platform, single-architecture scope.

What would change my assessment is evidence of broader reach — CUDA support, coverage of widely used architectures beyond Qwen3.5, and independent benchmarks showing generation speed genuinely comparable to llama.cpp on the same hardware. If those arrive in the next few release cycles, this becomes a meaningful consolidation of the local AI tooling landscape rather than a convenience feature.

Key Questions

Which hardware does the new GGUF support work on?

Currently, an Apple Silicon Mac is required. You also need a recent PyTorch release supported by the published ggml-quantization kernel builds and the latest transformers from the main branch.

How do I load a GGUF model in transformers?

Pass the Hub model_id and the GGUF filename as gguf_file to from_pretrained for both the tokenizer and model. No extra configuration is needed; everything after that uses the standard transformers API.

Which quantization level should I choose?

Hugging Face suggests starting with Q4_K_M and trying Q5_K_M or Q6_K if you have more memory. The quality tradeoff of more aggressive quantization varies by model and task, so test on your actual workload.

Can I serve GGUF models with an OpenAI-compatible API?

Yes. transformers serve accepts a GGUF checkpoint in the form <model_id>:<filename>.gguf and exposes an OpenAI-compatible endpoint that clients like Jan or Pi can connect to.

Does this replace llama.cpp, Ollama, or LM Studio?

Hugging Face does not claim parity beyond performance “close to llama.cpp” on Apple Silicon, using llama.cpp’s own ggml kernels. llama.cpp-derived tools remain standalone options; the new feature mainly benefits developers already working in the transformers and PyTorch ecosystem.

Source: Hugging Face

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

SenseTime Releases SenseNova U1 Pro Image Model With Up To 8K Output – TechNode

SenseTime has released SenseNova U1 Pro, an image generation model supporting output resolutions up to 8K, according to a company announcement.

Better Prompt Caching For GPT-6

OpenAI has published a headline about better prompt caching for GPT-6, but no article details are available to confirm what changed.

Bring Your Spreadsheet Data To Life With Sheets Canvas

Sheets canvas uses Gemini to turn spreadsheet data into interactive dashboards and trackers that stay synced with Google Sheets.

Higgsfield AI Ships New Video Features In A Day With GPT-6 Astra

OpenAI says Higgsfield AI used GPT-6 Astra to ship new video features in a single day. What is confirmed, what is claimed, and what it means.