Research library / Whitepaper

Real-Time Fraud Scoring at Wire Speed

In-Process AI Inference on Payment Event Streams — No Model Server. No GPU. No Spark.

Avg throughput
126,224 records · 5.14 min
409.4 EPS
Peak throughput
sustained · no warmup gap
511 EPS
Inference avg
ONNX in-process · CPU only
6.8 ms
vs baseline
same JAR · same config
8.12×
Dropped records
across all 5 benchmark runs
0
GC overhead
33 ms max pause · G1GC
0.049%
Provenance headers
per output record · SHA-256
22
Live heap (post-GC)
stable · load-independent
55 MB

Tuning progression · 5 runs · same hardware

RunAvg EPSPeak EPSInfer avgKey change
1 · Baseline50.476.918.7 mspool=1 · no batching · batch=32
2 · Thread tuning52.3102.677.3 mspool=4 · intra=3 · thread collision
3 · Batching on145.6294.416.4 msbatching enabled · pool=4 · intra=2
4 · Batch aligned280.0419.911.1 mspipeline batch=16 → matches embed max
5 · This run ✓409.4511.16.8 msJVM fully JIT-warm · identical config

Pipeline spec

Transform chain: STRING_TO_WIREEVENT → DJL_EMBEDDING → FRAUD_SCORE → KAFKA
Model: all-MiniLM-L6-v2 (ONNX) · champion / production
Hardware: Intel i9-8950HK · CPU only · no GPU
Config: parallelism=4 · pool=4 · intra-op=2 · batch=16
Kafka: 12 partitions · lz4 · 256 KB batch
GC: G1GC · 4 GB heap · 3 threads · MaxPause=50ms

Author

Steven Lopez
Solutions Architect & Creator
StreamKernel LLC
steven.lopez@streamkernel.io
linkedin.com/in/steve-lopez-b9941
May 20, 2026 · v0.2.0

01 Executive summary

Financial institutions process tens of millions of payment events daily. Every millisecond of scoring latency increases fraud exposure. Every external network hop to an inference service is a failure surface, a compliance boundary, and a cost center. Every scored transaction that loses its audit trail creates regulatory risk.

StreamKernel eliminates all three problems simultaneously. It is a Java 21 event processing kernel that executes AI inference in-process — inside the same JVM as the pipeline, with no external model-serving infrastructure required. Fraud scores, embeddings, decisions, and full cryptographic provenance travel together in every output event, from the moment a transaction enters the pipeline to the moment it lands in Kafka, MongoDB, Delta Lake, or any downstream system.

In a live 5-minute benchmark run on May 20, 2026:
→ 126,224 synthetic payment transactions scored end-to-end
→ Zero record loss · Zero errors · Zero DLQ events
→ Every output record carried 22 cryptographic provenance headers
→ GC overhead: 0.049% · Heap at close: 1,214 MB of 4 GB (28.3%)
→ Hardware: Intel i9-8950HK laptop — CPU only, no GPU

This white paper presents the full technical profile of StreamKernel's financial fraud scoring pipeline: architecture, transform chain, benchmark methodology, raw metrics, and fraud decisioning output. The results establish a credible performance and integrity floor for production deployment in regulated financial services environments.

02 Problem statement

The Latency and Auditability Gap in Real-Time Fraud Decisioning

Modern fraud detection has converged on machine learning as the primary scoring mechanism. Gradient-boosted trees, neural networks, and embedding-based semantic similarity models now underpin the risk engines of the world's largest card networks, banks, and payment processors. The models themselves are no longer the bottleneck. The infrastructure that serves them is.

The External Inference Service Problem

The dominant deployment pattern for ML-powered fraud scoring today is the external inference service: a model server (TensorFlow Serving, Triton, SageMaker, or a cloud endpoint) that receives a serialized feature vector over HTTP or gRPC, runs inference, and returns a score. This pattern has three compounding problems:

  • Network latency adds 5–50 ms per transaction, every transaction. At 10,000 TPS, a 10 ms average round-trip to an inference service adds 100 seconds of cumulative latency per second of transactions.
  • The inference service is a new failure domain. It must be deployed, scaled, monitored, secured, and versioned independently. Every model update is a multi-system coordination event.
  • Provenance is fragmented. The score arrives back at the pipeline as a floating-point number. The model version, feature set, inference timestamp, and configuration that produced it live in separate systems — if they are captured at all.

The Regulatory Dimension

For institutions subject to SR 11-7, DORA, PSD2 Article 95, or equivalent model risk management frameworks, the provenance gap is not an operational inconvenience — it is a compliance requirement. Regulators expect institutions to demonstrate, for any model-driven decision, which model version produced the score, on what inputs, at what time, under what configuration. Assembling that audit trail from logs and separate model management systems after the fact is error-prone and expensive.

What Is Required

A fraud scoring architecture that is production-ready for regulated financial services must satisfy four properties simultaneously:

  • Low and deterministic latency — sub-100 ms end-to-end from event ingestion to scored output in Kafka or equivalent
  • Zero record loss — every transaction that enters the pipeline must produce an output event, or a guaranteed dead-letter record
  • Per-event audit provenance — the model version, feature version, configuration hash, inference timestamp, and decision rationale must travel with every output event
  • Transport independence — the pipeline must be deployable on Kafka, Pulsar, MongoDB, Delta Lake, or bare-metal message buses without rewriting the scoring logic

StreamKernel was designed to satisfy all four properties in a single deployable artifact.

03 Technical architecture

StreamKernel Architecture

StreamKernel is a transport-agnostic event pipeline kernel. Its core runtime is a Java 21 process that orchestrates a configurable transform chain between a pluggable source and one or more pluggable sinks. The design principle is infrastructure collapse: every component that would normally be a separate service — the model server, the feature store cache, the schema registry, the audit log — runs inside the same JVM as the pipeline.

The Transform Chain

Every StreamKernel pipeline is defined by a transform chain: a named, ordered sequence of transformer plugins applied to each event batch. For financial fraud scoring, the chain is:

StepPluginFunction
1STRING_TO_WIREEVENTParses raw payment event text into a typed WireEvent envelope with transaction ID, amounts, merchant, channel, and country fields extracted
2DJL_EMBEDDINGRuns all-MiniLM-L6-v2 (ONNX) in-process via Deep Java Library to produce a 384-dimension normalized embedding vector for each payment event
3FRAUD_SCOREApplies a deterministic scoring model to the embedding vector, computing a continuous fraud score [0.0–1.0] and assigning a risk band and decision with reason codes
4KAFKA (sink)Writes the enriched, scored event to a Kafka topic with full provenance headers attached at the producer level

In-Process AI Inference

The DJL_EMBEDDING transformer is the architectural centerpiece of the fraud scoring pipeline. Rather than dispatching to an external model server, StreamKernel loads the ONNX model file directly into the JVM at startup using Deep Java Library and the ONNX Runtime engine (v1.20.0). The model remains resident in memory for the pipeline's lifetime.

Each transaction's text payload is tokenized using the Hugging Face tokenizer (max 16 tokens), passed through the ONNX runtime, and the resulting embedding vector is normalized and attached to the WireEvent. The entire inference operation — tokenization, ONNX forward pass, normalization — runs on the same thread that owns the event batch, with no cross-process communication.

In-process inference means: no serialization, no network round-trip, no external service dependency.
The model is a local file. The ONNX runtime is a JVM library. Inference is a method call.
This is what 'infrastructure collapse' means in practice.

MLflow Model Lifecycle

StreamKernel integrates with MLflow for model identity and lifecycle management. The champion model is referenced by alias (champion) and stage (production), and its identity is cryptographically anchored to every output event via a SHA-256 model reference hash. Model promotion and rollback can be executed at runtime without pipeline restarts — the HotSwappableDjlEmbeddingTransformer enables live model replacement with zero record loss.

Model name: streamkernel-minilm-onnx
Active alias: champion
Active stage: production
Model ref SHA-256: 12f054f7c8e13bae3db9b20761dd44e5c2c598749ae48b0f5a6ce0d7d5f7b16d

Transport-Agnostic Sink Architecture

The same fraud scoring transform chain — STRING_TO_WIREEVENT → DJL_EMBEDDING → FRAUD_SCORE — runs unchanged regardless of the output destination. The sink is a pluggable component loaded at startup from the plugin catalog. In the benchmark configuration, the sink is Kafka with at-least-once producer settings (acks=1, enable.idempotence=false, retries=MAX_INT). The same chain can be redirected to MongoDB (vector or document sink), Delta Lake, Pulsar, Snowflake, or PostgreSQL by changing a single configuration property.

Available sinks in the benchmark build's plugin catalog:

  • KAFKA — producer with configurable acks, lz4 compression, idempotent mode available
  • MONGO_INSERT — document sink with WiredTiger compression
  • MONGO_VECTOR — vector embedding sink for similarity search
  • DELTA — Delta Lake sink for analytical workloads
  • PULSAR — Apache Pulsar with configurable schema registry
  • POSTGRES / PGVECTOR — relational and vector storage
  • DEVNULL — no-op sink for benchmarking transform chain in isolation

Security and Compliance Architecture

StreamKernel's security layer is plugin-based. The benchmark run uses PERMIT_ALL (no authentication) for a clean throughput baseline. Production deployments can substitute OPA_SIDECAR for Open Policy Agent-based per-batch authorization, or mTLS+OPA for mutual TLS with policy evaluation. Authorization refresh suppression, fail-closed caching, and configurable token refresh cycles are available in the security plugin interface.

04 Benchmark

Benchmark Methodology and Configuration

Test Configuration

The benchmark was executed using StreamKernel's automated test runner against a single pipeline configuration. All parameters were fixed for the duration of the run. No tuning was performed between the start and stop signals.

Test name: streamkernel_financial_fraud_scoring_5mRun ID: run-financial-fraud-01
Duration: 5.14 minutes (300 s target)Heap: 4 GB (G1GC, 3 threads, MaxGCPauseMillis=50)
Parallelism: 4 worker threadsBatch size: 16 records
In-flight limit: 768 recordsExecutor mode: FIXED
Embedding pool size: 4 predictors (batching enabled, max.size=16)Tokenizer max length: 16 tokens
Cache: DISABLED (NOOP)Source: SYNTHETIC PAYMENTS profile, 384-char payload
Max records/sec: Unlimited (no source-side throttle)Kafka partitions: 12
Kafka acks: 1Kafka compression: lz4

Hardware

Intel Core i9-8950HK — 6 cores / 12 threads — CPU only, no GPU, no accelerator
All inference is CPU-bound ONNX Runtime with 2 intra-op threads, 1 inter-op thread
These results represent a conservative performance floor, not a ceiling.
GPU-accelerated or multi-core pooled configurations will materially outperform these numbers.

Measurement Approach

Metrics were collected via StreamKernel's embedded Prometheus endpoint (port 8080) and captured to a .prom snapshot at run completion. The speedometer logs a 5-second rolling window BENCH line to stdout every 5 seconds throughout the run, providing time-series throughput and latency data independent of the Prometheus snapshot. The post-run Kafka record count was verified by consuming all partitions and comparing against the pipeline_in_total, pipeline_out_total, and pipeline_processed_total counters.

A config.sha256 (b10e981ee98c1d470b00c0390859186bb0e69ead5056504a497631daa2d79027) and model.ref.sha256 were computed at pipeline startup and attached to every Kafka output record, establishing a cryptographic link between the benchmark evidence and the exact code and model artifact used.

05 Results

Benchmark Results

Headline Numbers

Records Processed
in = out = processed
126,224
Avg EPS
events per second
409.4
Record Loss
drops · DLQ · errors
0
GC Overhead
G1GC, 3 threads
0.049%
Heap at Close
28.3% of 4 GB
1,214 MB
Inference Avg
ONNX, CPU, batched
6.8 ms
Max GC Pause
target ≤ 50 ms
33 ms
Provenance Headers
per output record
22

Data Integrity

pipeline_in_total = pipeline_out_total = pipeline_processed_total = 126,224
pipeline_dropped_total = 0
pipeline_dlq_total = 0 · pipeline_dlq_errors_total = 0
pipeline_source_errors_total = 0 · pipeline_auth_errors_total = 0
pipeline_denied_total = 0
kafka_sink_sent_ok_total = 126,224 (matches pipeline_out_total exactly)

Perfect record integrity was maintained for the full 5.14-minute run. Every transaction that entered the pipeline produced a scored output record in Kafka. No records were dropped due to backpressure, timeout, serialization failure, or model error.

Per-Stage Latency

StageAvg LatencyMax LatencyCount
Tokenize (HF tokenizer)1.15 ms11 ms15,617
ONNX inference (all-MiniLM-L6-v2)6.8 ms52 ms126,224
WireEvent string encode1.00 ms1 ms126,224
Kafka producer send13.8 ms50 ms126,224
Kafka request latency (broker RTT)3.9 ms—126,224
Kafka record queue time11.0 ms—126,224

The ONNX inference stage averages 6.8 ms per record — the best result across the entire tuning series. This is achieved with a pool of 4 ONNX predictors, each running batches of 16 records, with 2 intra-op threads per predictor. The batching engine maintained 100% fill rate throughout the run, meaning no predictor ever fired an under-filled batch.

Throughput Profile

The pipeline sustained a throughput band of 118–511 EPS across the 5-minute run, measured in 5-second rolling windows. The pipeline was fully saturated from the first window — no warm-up gap. GC-related dips were brief and transient; throughput recovered within one 5-second window every time. The pipeline never stalled, shed records, or entered backpressure hold.

WindowEPSWindowEPSWindowEPS
0:05357.71:45367.63:25256.0
0:10511.11:50330.23:30354.9
0:15390.41:55306.63:35333.4
0:20383.62:00262.23:40281.2
0:25297.52:05118.63:45284.7
0:30246.12:10271.73:50355.8
0:35269.42:15265.73:55319.2
0:40297.72:20188.54:00275.3
0:4557.62:2551.14:0532.0
0:50371.32:30294.24:10223.8
0:55364.82:35343.24:15134.2
1:00469.52:40470.14:20469.2

JVM and GC Health

G1GC configuration targeted a 50 ms max pause time with 3 GC threads. The 33 ms max pause is well within target and reflects the higher allocation rate at 409 EPS — proportional to the 7.47 GB total allocated over the run.

MetricValue
GC overhead at run end0.049%
G1 Young Generation pauses (total)18 (1 metadata threshold, 17 evacuation)
Total GC pause time291 ms over 5.14 minutes
Max single GC pause33 ms (target: ≤ 50 ms)
Heap used at run end1,214 MB of 4 GB (28.3%)
Live data size (post-GC)55.0 MB
Total heap allocated (run lifetime)7.47 GB
Heap occupancy after last GC2.95% of max
Live threads at close11 (peak: 17)

Kafka Sink Health

The Kafka producer operated across 12 topic partitions with lz4 compression and a 256 KB batch size. Kafka send averaged 13.8 ms with a 50 ms max — a significant improvement over earlier tuning runs. Broker RTT averaged 3.9 ms, returning to near-baseline levels as higher throughput amortized the partition coordination overhead.

Kafka MetricValue
Records sent OK126,224
Kafka request latency (avg)3.9 ms
Kafka record queue time (avg)11.0 ms
Kafka producer send time (avg)13.8 ms (max: 50 ms)
Producer buffer available (end)256 MB (fully available)
Topic partitions12
Compressionlz4
Producer modeacks=1, retries=MAX_INT

Note on partition distribution:
Partition distribution across 12 partitions was well-balanced at the throughput achieved. Higher throughput naturally smooths adaptive partitioning skew. Kafka's adaptive partitioning is enabled; this skew may be transient, driven by key distribution across the synthetic transaction IDs, or an artifact of the 12-partition topology relative to the 4-worker pipeline. Production deployments should monitor partition lag across all consumers and adjust the partitioning strategy (key-based or round-robin) to match downstream consumer group topology.

06 Audit provenance

Per-Event Provenance and Audit Trail

Every record written to the streamkernel-financial-fraud-scored Kafka topic carries 22 provenance headers in addition to the JSON payload. These headers are set by the pipeline at the producer level — they are not derived from the payload and cannot be modified by downstream consumers. They travel with the record through any Kafka consumer, stream processor, or sink that preserves headers.

Complete Provenance Header Manifest

Header KeyExample ValuePurpose
streamkernel.provenance.pipeline.idsk-financial-fraud-scoringPipeline identity
streamkernel.provenance.run.idrun-financial-fraud-01Run traceability
streamkernel.provenance.config.sha256735d81f5c727…e84f39Config SHA-256 fingerprint
streamkernel.provenance.model.namestreamkernel-minilm-onnxMLflow model name
streamkernel.provenance.model.aliaschampionActive MLflow alias
streamkernel.provenance.model.stageproductionMLflow stage at inference
streamkernel.provenance.model.ref.sha25612f054f7c8e1…7b16dModel artifact SHA-256
streamkernel.provenance.inference.timestamp2026-05-19T15:07:17.748ZNanosecond timestamp
streamkernel.provenance.transform.chainSTRING_TO_WIREEVENT, DJL_EMBEDDING,FRAUD_SCOREFull transform chain applied
streamkernel.provenance.transform.versionpayment-fraud-score-v1Named transform version
streamkernel.provenance.feature.versionpayment-risk-features-v1Feature engineering version
streamkernel.provenance.prompt.versionnot-applicableLLM prompt version (N/A)
streamkernel.provenance.source.typeSYNTHETICSource plugin used
streamkernel.provenance.sink.typeKAFKASink plugin used
streamkernel.provenance.sink.authPLAINTEXTTransport auth mode
streamkernel.provenance.security.typePERMIT_ALLAuthorization policy
fraud.model.versiondeterministic-fraud-score-v1Fraud model version
fraud.score0.3700Computed fraud score
fraud.risk_bandLOW / MEDIUM / HIGHRisk classification
fraud.decisionAPPROVE / WATCH / REVIEWFinal decision
fraud.reason_codesCROSS_BORDER, NEW_DEVICE…Decision rationale codes
fraud.review.threshold0.7200Review threshold used

The config.sha256 and model.ref.sha256 headers provide a cryptographic chain of custody from every output event back to the exact pipeline configuration and model artifact that produced it. For model risk management under SR 11-7 or equivalent frameworks, this means the full audit trail for any scored transaction is available by inspecting the Kafka record — no external log join required.

07 Output examples

Sample Fraud Scoring Output

The following records were sampled from the live Kafka topic immediately after the benchmark run. They represent the three risk tiers produced by the FRAUD_SCORE transform: APPROVE (score < 0.45), WATCH (0.45 ≤ score < 0.72), and REVIEW (score ≥ 0.72).

Record 1 — APPROVE / LOW Risk

FieldValue
Transaction IDTXN-100192
AccountACCT-15952
AmountUSD 275.04
Merchant / Channelgrocery / bill_pay
CountryAE (United Arab Emirates)
Fraud Score0.37
Risk BandLOW
DecisionAPPROVE
Reason CodesCROSS_BORDER, NEW_DEVICE

Record 2 — WATCH / MEDIUM Risk

FieldValue
Transaction IDTXN-100193
AccountACCT-15983
AmountUSD 276.41
Merchant / Channelfuel / card_present
CountrySG (Singapore)
Fraud Score0.61
Risk BandMEDIUM
DecisionWATCH
Reason CodesCROSS_BORDER, NEW_DEVICE, HIGH_VELOCITY

Record 3 — REVIEW / HIGH Risk

FieldValue
Transaction IDTXN-100194
AccountACCT-16014
AmountUSD 277.78
Merchant / Channeltravel / card_not_present
CountryBR (Brazil)
Fraud Score0.99
Risk BandHIGH
DecisionREVIEW
Reason CodesCROSS_BORDER, UNKNOWN_DEVICE, ELEVATED_VELOCITY, SANCTIONS_SCREEN_HIT, HIGHER_RISK_CHANNEL

The FRAUD_SCORE transform applies five independent risk signals — velocity, device trust, channel risk, sanctions screening, and cross-border exposure — and aggregates them into a continuous [0.0–1.0] score.
The two configurable thresholds (WATCH ≥ 0.45, REVIEW ≥ 0.72) are externalized in the pipeline configuration and captured in the fraud.review.threshold provenance header on every record, so downstream systems can reconstruct the decisioning logic without access to the pipeline source.

08 Use cases

Financial Services Deployment Scenarios

StreamKernel's architecture is particularly well-suited to financial services workloads where the combination of low latency, regulatory auditability, and transport flexibility is non-negotiable.

Card Network Real-Time Authorization

Card authorization networks process thousands of transactions per second with sub-100 ms SLAs on scoring decisions. StreamKernel's in-process inference eliminates the external model-serving hop, reducing the scoring path to tokenization + ONNX inference + rule evaluation — all within the same JVM. The pipeline can be deployed as a sidecar to the authorization engine or as a standalone scoring service with a Kafka input/output interface.

Bank Transaction Monitoring

Anti-money laundering and fraud monitoring systems at retail banks consume transaction feeds from core banking systems and apply behavioral and semantic scoring to detect anomalous patterns. StreamKernel can consume these feeds from Kafka, Pulsar, or REST endpoints, apply embedding-based semantic similarity scoring, and write enriched events to Delta Lake or MongoDB for downstream analytics — all with per-event audit headers that satisfy model risk management documentation requirements.

Payment Processor Chargeback Prediction

Payment processors seeking to predict chargeback risk at authorization time can deploy StreamKernel as an enrichment layer between the payment gateway and the downstream risk engine. The embedding vector produced by DJL_EMBEDDING can feed both the real-time scoring model and a vector store (MONGO_VECTOR or PGVECTOR) for post-hoc similarity analysis against historical chargeback patterns.

Fintech and Neo-Bank Deployment

For fintechs operating without legacy infrastructure, StreamKernel's transport-agnostic architecture means the entire fraud scoring pipeline can be deployed against a managed Kafka service (Confluent Cloud, MSK, Aiven) or Pulsar without on-premises infrastructure. The MLflow integration provides model governance without requiring a dedicated model platform — a single MLflow tracking server is sufficient.

Air-Gapped and High-Security Environments

For institutions operating in air-gapped or highly restricted network environments — including defense-adjacent financial services, sovereign wealth funds, and central banks — StreamKernel's fully in-process architecture is a significant operational advantage. The model, tokenizer, and scoring logic are packaged in a single fat JAR. There are no outbound calls to model serving infrastructure, no cloud dependencies, and no runtime model downloads. The pipeline operates as a standalone process with local file access only.

09 Conclusion

Conclusion

Real-time fraud scoring at production scale requires an architecture that resolves three tensions simultaneously: inference speed versus infrastructure complexity, throughput versus auditability, and flexibility versus operational overhead. External model serving infrastructure solves none of these tensions — it trades inference latency for operational complexity, and sacrifices auditability for architectural separation.

StreamKernel resolves all three by collapsing the inference infrastructure into the pipeline. The benchmark results presented in this paper demonstrate that a CPU-only laptop running StreamKernel can process 126,224 payment transactions in 5 minutes at 409 EPS average — with zero record loss, 6.8 ms average ONNX inference latency, 100% batch fill rate, and full 22-field cryptographic provenance on every output event.

These are conservative numbers. Production hardware — multi-socket servers, higher core counts, expanded embedding pools, GPU ONNX execution providers — will produce proportionally higher throughput with equivalent or better integrity guarantees.

Benchmark Summary

126,224 records · 5.14 min · 409.4 EPS avg · 511 EPS peak
Zero record loss · Zero errors · Zero DLQ
0.049% GC · 1,214 MB heap · 33 ms max pause

22 provenance headers per event
Cryptographic config + model SHA-256
MLflow model lifecycle integration
Transport-agnostic · CPU-only deployable

For licensing, partnership, or enterprise deployment inquiries: steven.lopez@streamkernel.io
LinkedIn: linkedin.com/in/steve-lopez-b9941

Commercial path

Want to turn this paper into a concrete evaluation?

Bring the action you have in mind and we will map it to what the runtime does today.