Research library / Article

MiniLM-L6-v2 on the JVM: How far can you push CPU inference?

Model: all-MiniLM-L6-v2 · Runtime: ONNX Runtime 1.20.0 / Java 21 · Hardware: 12-logical-core CPU, Windows 11, High Performance · Date: April 14, 2026

Why this benchmark exists

StreamKernel is a JVM-native real-time AI enrichment engine — a single process that ingests streaming records, runs ONNX inference in-process, and writes enriched output to a sink. No Python. No sidecar. No network hop to a model server.

That architectural bet — inference lives inside the JVM, not beside it — is the core value proposition for regulated and air-gapped environments. But it raises an obvious question: how fast is fast enough on CPU, and which configuration choices actually move the needle?

To answer that precisely, we built djl-onnx-bench: a standalone Java benchmark harness that calls ORT directly, with controlled warmup, cooldown between runs, and a sweep runner that randomizes execution order to prevent thermal accumulation bias. This post reports the results of an 8-configuration sweep.

"The goal was not to find the fastest number we could publish. It was to understand the shape of the performance surface — where INT8 wins, where thread count matters, and where the tail latency tells a different story than throughput."

StreamKernel Engineering

What we tested

Two model variants — fp32 (standard ONNX export) and INT8 (dynamically quantized) — across two sequence lengths (s16, s32) and two intra-op thread counts (3, 4), all at batch size 32. Eight configurations total, randomized execution order, 10 warmup iterations discarded, 40 measured iterations per run, 10-second cooldown between configurations.

Methodology note

Sequence lengths of 16 and 32 tokens are shorter than StreamKernel's production tokenizer.max.length=128. These values were chosen to map the performance surface efficiently. Inference scales roughly O(n²) with sequence length in transformer attention, so production numbers at seq=128 will be proportionally slower — but the relative rankings between configurations hold.

All runs on High Performance Windows power plan, plugged in, CPU core parking disabled, PCI-E ASPM off.

Peak EPS
347
Configurations
8
INT8 / FP32 range
2.9×

Results matrix

Sorted by EPS — what actually won

INT8 quantized · FP32 standard

ModelSeqThreadsAvg msP50 msP95 msP99 msEPSP95/P50
1int816391.9782.05133.51192.31347.821.63×
2int816497.1587.23147.15196.72329.281.69×
3fp32164100.0594.57137.51158.31319.761.45×
4fp32163120.31111.82172.38188.28265.921.54×
5int8324149.64142.90238.77290.61213.811.67×
6int8323187.55164.83258.13306.04170.601.57×
7fp32324245.14234.65329.53426.50130.521.40×
8fp32323265.84253.88428.02582.37120.361.69×

What the data says

Finding 01

INT8 wins on throughput — but not on tail consistency

The top EPS belongs to int8 s16 intra3 at 347.82. But fp32 s16 intra4 has the tightest tail of the sequence-16 runs: P95/P50 of 1.45×. For streaming pipelines where tail latency causes queue buildup, that distribution shape may matter more than raw throughput.

Finding 02

INT8's advantage grows with sequence length

At s16, int8 intra4 is only 3% faster than fp32 intra4 (97ms vs 100ms). At s32, that gap widens to 39% (150ms vs 245ms). INT8 quantization reduces memory bandwidth pressure — and longer sequences have larger attention matrices where that pressure compounds.

Finding 03

Thread count has a crossover point

intra3 beats intra4 for INT8 at short sequences (347 vs 329 EPS), but loses at longer ones (171 vs 214 EPS at s32). At s16 the work is small enough that 3 threads saturate it without coordination overhead. At s32, the fourth thread earns its keep.

Finding 04

Run 6 of 8 shows thermal accumulation

The fp32 s32 intra3 run — sixth of eight, after four consecutive CPU-intensive benchmarks — has a P99/P50 ratio of 2.29×. The 10-second cooldown between runs was insufficient. This is why randomized execution order and per-run cooldowns are non-negotiable in CPU benchmarking.

The CPU budget constraint

These results were collected on a 12-logical-core machine. The winning configurations aren't random — they respect a hard constraint: total ONNX threads ≤ physical cores. Violating this boundary, as we confirmed in separate StreamKernel pipeline runs, produces severe contention: a pool.size=3, intra=6 configuration (18 threads on 12 cores) regressed from 410ms to 1,460ms average batch time.

The sweet spot for this hardware is 2–3 predictors × 4–6 intra-op threads = 8–12 total threads. Beyond that, you're not buying parallelism — you're buying scheduling contention.

For StreamKernel users

The INT8 model is a drop-in replacement. Set ai.embedding.model.uri to your model.int8.onnx path. No other changes required. Based on this sweep, expect 35–50% lower batch latency at production sequence lengths compared to fp32, with proportionally higher end-to-end pipeline throughput.

Benchmark methodology

Harness
djl-onnx-bench (Java 21)
Runtime
ORT 1.20.0 / DJL 0.32.0
Model
all-MiniLM-L6-v2 (fp32 + INT8)
Batch size
32 (fixed)
Warmup
10 iterations (discarded)
Measured
40 iterations per config
Execution order
Randomized
Cooldown
10s between configs
Power plan
High Performance (AC)
Input type
Synthetic pre-tokenized tensors

Input preparation averaged 0.001–0.002ms per batch — effectively zero. All reported latency is pure ORT inference time. Tokenization cost, DJL pooling, and sink write latency are measured separately in full StreamKernel pipeline benchmarks.

What's next

The immediate next step is running the INT8 model through the full StreamKernel pipeline at seq=128 — the production tokenizer length — to measure end-to-end throughput uplift. Based on the scaling observed between s16 and s32 in this sweep, INT8 at seq=128 should cut batch latency by 35–50% relative to the fp32 baseline, pushing pipeline throughput well above the current 18.8 records/sec ceiling established in Run 3.

Longer term, the same benchmark harness will be used to evaluate alternative embedding models and quantization strategies as StreamKernel's model zoo expands. The sweep infrastructure supports arbitrary config matrices — add a row to sweep.csv and the runner handles the rest.

StreamKernel
Real-time AI enrichment for regulated infrastructure.
http://medium.com/@lopezstevie

Commercial path

Want to turn this paper into a concrete evaluation?

Bring the action you have in mind and we will map it to what the runtime does today.