APPLE SILICON · PRIVACY-FIRST · ON-DEVICE

Edge LLM inference,
measured on real hardware.

Real benchmark data from Apple Silicon — tokens/sec, energy efficiency, memory bandwidth. Every byte stays on-device. No cloud. No logs. No surveillance.

framework: mlx-lm 0.12.1 device: Mac14,2 Apple M2 run: 2026-06-05

peak speed

39.7

tokens / second

first token

255ms

time to first byte

efficiency

2.6

tokens / joule

memory used

1.1GB

of 8 GB

models tested

2

open-weight models

data leaves device

0

privacy guaranteed

Model benchmarks

Mac14,2 Apple M2 · macOS 15.5 · click column headers to sort

mlx-lm 0.12.1
MODELQUANTTOK/S LATENCY RAM TOK/J MMLU SCORE
Llama 3.2 1B
1.2B params · General purpose
Q4101.11000.96.7
67
Llama 3.2 3B
3.2B params · General purpose
Q439.72551.12.6
26

Each model ran 1 warmup passes then 5 timed inference passes on a fixed 35-token prompt. Energy estimated from powermetrics.

Performance charts

TOKENS / SECOND · M2

MEMORY BANDWIDTH OVER TIME

EFFICIENCY · TOK / JOULE BY CHIP

Zero data egress

Every inference runs entirely on-chip. Prompts, completions, and intermediate activations never touch a network.

Unified memory advantage

Apple's shared CPU/GPU memory pool eliminates costly data transfers, cutting latency and energy use versus discrete GPU setups.

Neural Engine acceleration

The M3 ANE handles quantized model operations at up to 18 TOPS, offloading the CPU and reducing thermal throttling.

Live inference demo

cloud baseline via groq api · same models as the benchmark · latency comparison shown

live inference · cloud baseline

// try a prompt

CLOUD VS LOCAL

This demo uses Groq's cloud API as a reference. The benchmark table shows the same models running locally on M2 — typically ~30% lower latency with zero data leaving the device.

METHODOLOGY

Inferencemlx-lm 0.12.1
EnergymacOS powermetrics
QualityMMLU 5-shot
Runs5 timed passes