AI infrastructure

Inference belongs in the event path.

AI enrichment should not require a sidecar, a model-serving hop, and custom audit glue for every record. StreamKernel keeps ONNX inference, model labels, health-aware promotion and rollback, provenance, and delivery inside the pipeline boundary.

In-process enrichment

DJL + ONNX, MLflow governance, provenance, and destination delivery in one runtime.

The commercial differentiator is the AI path: model-aware operational movement without splitting every record across separate services and audit systems.

  • Inline ONNX inference before delivery.
  • MLflow registry support for live promotion, health checks, automated rollback, and blocked re-promotion evidence.
  • Model/version provenance on enriched output.
  • Kafka, MongoDB Vector, Postgres, PostgreSQL pgvector, Delta Lake, Snowflake, and custom sink paths.
  • OpenTelemetry source intake and local MCP tools for operational AI review.
In-process AI enrichment path
StreamKernel JVM Events Kafka / OTel Policy fail closed ONNX DJL inference model labels Destinations Kafka, MongoDB pgvector, Delta MLflow registry

AI use cases

Operational enrichment paths the website should make easy to understand.

Kafka -> ONNX -> Kafka

Enrich events inline and republish without a sidecar model-serving hop.

Kafka -> ONNX -> MongoDB Vector

Generate embeddings and insert vector records from the same pipeline boundary.

Kafka -> ONNX -> PostgreSQL pgvector

Score or embed records in-process and write vector-ready rows to PostgreSQL pgvector targets.

Any source -> ONNX -> Postgres

Deliver scored records from Kafka, Pulsar, REST, or custom sources into a relational store.

Pulsar -> ONNX -> Delta Lake

Run transport-agnostic AI enrichment into a lakehouse destination.

MLflow -> Delta Lake

Bootstrap model artifacts from the registry and write enriched records.

Live model swap

Promote MLflow models during a run, observe health, and fall back automatically when a fresh version degrades.

OpenTelemetry -> policy -> sink

Accept OTLP logs, metrics, or traces and send them through the same policy, redaction, enrichment, and DLQ path.

MCP agent control

Expose local agent-readable status, config validation, dependency health, benchmark summaries, and guarded model lifecycle tools.

Vertical inference

The inference types buyers recognize before they care about the pipeline.

The same in-process ONNX path can serve fraud scoring, clinical classification, sensor anomaly detection, and embedding generation without sending records through an external model hop.

Fraud scoring

Transaction events can receive model scores before they leave controlled financial systems.

Clinical classification

HL7/FHIR and telemetry streams can be classified by a local model without sending PHI to an external inference endpoint.

Sensor anomaly detection

Defense and industrial sensor streams can be scored at the edge in disconnected deployments.

Embedding generation

Operational records can be converted into vector-search-ready output for MongoDB Vector or PostgreSQL pgvector.

Inference provenance

Public-sector inference requests can carry model identity, input and output hashes, route decisions, and audit fingerprints.

Agent tool governance

Agent tool calls can be normalized, scored, routed, and audited before they affect enterprise systems.

AI proof strip

Evidence for the AI path starts with published local baselines.

The PRD calls for clear caveats, exact profiles, and reproducible evidence rather than a universal performance promise.

Governed fraud/AML pre-ingest screening

99,102 records/sec Apple M5 Pro · July 30–31, 2026 · 3-run mean · 10-minute runs · range 98,287–99,927 records/sec · CV 0.827% At-least-once; mTLS + Keycloak OIDC; fail-closed OPA; SHA-256 provenance on every record Current

Governed Kafka egress path (mTLS + OIDC + fail-closed OPA)

2.184M events/sec Apple M5 Pro · August 15, 2026 · 3-run mean · 10-minute runs · range 2.086M–2.354M events/sec · CV 6.766% acks=all with idempotence; Kafka sink mTLS + Keycloak OIDC; fail-closed OPA as a cached admission gate Current · cite range

Kafka WireEvent, at-least-once

2.201M records/sec Apple M5 Pro · July 28, 2026 · 3-run mean · 10-minute runs · range 2.161M–2.241M records/sec · CV 1.833% At-least-once, acks=1, no governance path Current

Kafka WireEvent, exactly-once

1.947M–2.312M records/sec observed range Apple M5 Pro · July 29, 2026 · 8-run mean · 10-minute runs · range 1.947M–2.312M records/sec (mean 2.061M) · CV 5.338% Exactly-once: acks=all, idempotence, one transaction commit per 100,000-record pipeline batch Current · cite range

Kafka-sink NOOP, source-paced capacity validation

1.995M records/sec Apple M5 Pro · July 30, 2026 · 5-run mean · 10-minute runs · range 1.984M–2.000M records/sec · CV 0.334% At-least-once, acks=1, no governance path Sustained capacity · not a ceiling

Kafka → MiniLM-L6-v2 ONNX CPU → Kafka, batching enabled

1,386 records/sec Apple M5 Pro · August 15, 2026 · 3-run mean · 10-minute runs · range 1,362.40–1,408.69 records/sec · CV 1.674% At-least-once; end-to-end processed records/sec, not isolated model operations/sec Current

Commercial path

Review the AI enrichment path against your models and destinations.

Bring the event source, model governance requirements, destination path, and audit expectations.