Kafka -> ONNX -> Kafka
Enrich events inline and republish without a sidecar model-serving hop.
AI infrastructure
AI enrichment should not require a sidecar, a model-serving hop, and custom audit glue for every record. StreamKernel keeps ONNX inference, model labels, health-aware promotion and rollback, provenance, and delivery inside the pipeline boundary.
In-process enrichment
The commercial differentiator is the AI path: model-aware operational movement without splitting every record across separate services and audit systems.
AI use cases
Enrich events inline and republish without a sidecar model-serving hop.
Generate embeddings and insert vector records from the same pipeline boundary.
Score or embed records in-process and write vector-ready rows to PostgreSQL pgvector targets.
Deliver scored records from Kafka, Pulsar, REST, or custom sources into a relational store.
Run transport-agnostic AI enrichment into a lakehouse destination.
Bootstrap model artifacts from the registry and write enriched records.
Promote MLflow models during a run, observe health, and fall back automatically when a fresh version degrades.
Accept OTLP logs, metrics, or traces and send them through the same policy, redaction, enrichment, and DLQ path.
Expose local agent-readable status, config validation, dependency health, benchmark summaries, and guarded model lifecycle tools.
Vertical inference
The same in-process ONNX path can serve fraud scoring, clinical classification, sensor anomaly detection, and embedding generation without sending records through an external model hop.
Transaction events can receive model scores before they leave controlled financial systems.
HL7/FHIR and telemetry streams can be classified by a local model without sending PHI to an external inference endpoint.
Defense and industrial sensor streams can be scored at the edge in disconnected deployments.
Operational records can be converted into vector-search-ready output for MongoDB Vector or PostgreSQL pgvector.
Public-sector inference requests can carry model identity, input and output hashes, route decisions, and audit fingerprints.
Agent tool calls can be normalized, scored, routed, and audited before they affect enterprise systems.
AI proof strip
The PRD calls for clear caveats, exact profiles, and reproducible evidence rather than a universal performance promise.
Governed fraud/AML pre-ingest screening
99,102 records/sec Apple M5 Pro · July 30–31, 2026 · 3-run mean · 10-minute runs · range 98,287–99,927 records/sec · CV 0.827% At-least-once; mTLS + Keycloak OIDC; fail-closed OPA; SHA-256 provenance on every record CurrentGoverned Kafka egress path (mTLS + OIDC + fail-closed OPA)
2.184M events/sec Apple M5 Pro · August 15, 2026 · 3-run mean · 10-minute runs · range 2.086M–2.354M events/sec · CV 6.766% acks=all with idempotence; Kafka sink mTLS + Keycloak OIDC; fail-closed OPA as a cached admission gate Current · cite rangeKafka WireEvent, at-least-once
2.201M records/sec Apple M5 Pro · July 28, 2026 · 3-run mean · 10-minute runs · range 2.161M–2.241M records/sec · CV 1.833% At-least-once, acks=1, no governance path CurrentKafka WireEvent, exactly-once
1.947M–2.312M records/sec observed range Apple M5 Pro · July 29, 2026 · 8-run mean · 10-minute runs · range 1.947M–2.312M records/sec (mean 2.061M) · CV 5.338% Exactly-once: acks=all, idempotence, one transaction commit per 100,000-record pipeline batch Current · cite rangeKafka-sink NOOP, source-paced capacity validation
1.995M records/sec Apple M5 Pro · July 30, 2026 · 5-run mean · 10-minute runs · range 1.984M–2.000M records/sec · CV 0.334% At-least-once, acks=1, no governance path Sustained capacity · not a ceilingKafka → MiniLM-L6-v2 ONNX CPU → Kafka, batching enabled
1,386 records/sec Apple M5 Pro · August 15, 2026 · 3-run mean · 10-minute runs · range 1,362.40–1,408.69 records/sec · CV 1.674% At-least-once; end-to-end processed records/sec, not isolated model operations/sec CurrentCommercial path
Bring the event source, model governance requirements, destination path, and audit expectations.