Research library / Visual story

No cloud. No calls home. No compromise.

Most AI inference platforms require internet connectivity, cloud APIs, or managed model registries. StreamKernel was built without that assumption.

The challenge

Regulated environments reject phone-home architectures

Defense, intelligence, healthcare, and critical infrastructure ops can't afford a model that dials out to validate a license or fetch weights at runtime.

Network isolation

SCIFs, classified networks, and OT environments have no egress. Period.

Data sovereignty

ITAR, FedRAMP, and HIPAA restrict where data can travel—or forbid it entirely.

Latency intolerance

Real-time sensor fusion and threat detection can't afford a round-trip to a remote API.

Disconnected ops

Ships, forward bases, and edge sites operate with intermittent or zero connectivity.

How it works

Everything runs in-process. No agents. No sidecars.

StreamKernel embeds ONNX/DJL AI inference directly inside the JVM pipeline runtime — no external model server, no gRPC call, no network hop.

Diagram comparing a standard cloud-dependent pipeline, which passes through a remote model server on the internet, with the StreamKernel approach, where the pipeline, ONNX inference, and MLflow registry all run inside one JVM.

Benchmark evidence

Real numbers from a sealed stack

Validated on a single-node JVM with zero cloud dependencies. These aren't lab conditions — this is the production-candidate configuration.

Records/sec
AI embedding throughput
336.8
Throughput improvement
from baseline run
28×
Record loss
at 200K+ docs / run
0
Docs/sec avg
raw MongoDB write
163K
Model swap cycle
promote → rollback
<30s
Entire runtime
no external deps
1 JAR

Air-gap readiness

Every requirement. Checked by design.

No outbound network calls at runtime

ONNX weights and MLflow model files are loaded from local disk or internal registry. Zero internet dependency post-deploy.

Single-JAR deployment artifact

One file. Copy it. Run it. No container registry pull, no cloud bootstrap, no license server ping.

Local MLflow model governance

Promote, demote, and roll back models against an air-gapped MLflow instance. Full lineage. Full audit trail. No Databricks managed cloud.

mTLS + OPA security profile

Validated at 366K ops/sec with TLSv1.3 and OPA policy enforcement active simultaneously. Hardened for classified transport.

Per-event model lineage stamping

Every enriched record carries metadata identifying the exact model version that produced it — auditability without a cloud telemetry pipeline.

Competitive landscape

The rest weren't built for dark environments

RequirementTypical ML platformsStreamKernel
Runtime connectivityRequired — license server, model hub, or telemetryNone — fully offline after deployment
Model governanceCloud-managed registry (SaaS dependency)Local MLflow — self-hosted, air-gapped
Inference locationRemote API call or sidecar processIn-process, within the JVM — no hop
Deployment artifactContainer images + orchestrator + registrySingle JAR, runs on bare JVM
Security postureVaries — TLS optional, policy not enforcedmTLS + OPA enforced at 366K ops/sec
Audit trailCloud-side logs onlyPer-event lineage in every output record

AI inference at stream speed. Inside the perimeter.

StreamKernel is the only JVM pipeline runtime with embedded ONNX/DJL inference and local MLflow governance — built for environments where the cloud is not an option.

steven.lopez@streamkernel.io

Commercial path

Want to turn this paper into a concrete evaluation?

Bring the action you have in mind and we will map it to what the runtime does today.