Back to Blog
Engineering

From Raw Events to Trust Score: How Karma3's Data Pipeline Works Under the Hood

Marcus Chen 14 min read
Data pipeline architecture from raw events to trust score

A behavioral trust score is the output of a pipeline. The inputs are raw session events: keystrokes, mouse movements, navigation transitions, touch events on mobile, form interactions. Between those raw inputs and the float between 0 and 1 that comes out of the score call, there is a sequence of processing steps with specific latency budgets and specific data contracts. This post walks through each stage.

We are not going to hand-wave about machine learning. The goal is to give engineering teams a concrete picture of what the pipeline does, where latency is incurred, and what the data looks like at each stage, because those questions come up repeatedly in integration evaluations.

Stage 1: Event collection and transmission

The SDK instruments the client environment and collects behavioral events into an in-memory event buffer. The event types are: keystroke events (character timing, not content), mouse trajectory samples at 100ms intervals, scroll depth and velocity, focus/blur events on form elements, touch events on mobile (pressure, duration, release velocity), and page navigation events with timing.

The event buffer has a size limit. On a typical login session from page load to form submit, the buffer accumulates approximately 400-1200 events depending on session length and user input activity. The buffer is not streamed in real time: it is held in memory and transmitted as a single encrypted payload at the moment the session action occurs (form submit, checkout proceed, etc.).

The transmission is a POST to our ingestion endpoint with the encrypted payload and the session ID in the header. The payload size is typically 4-12KB depending on session length. Transmission adds 8-20ms depending on geographic distance to the nearest ingestion point. We run ingestion infrastructure in four regions; SDK initialization resolves to the nearest one.

Stage 2: Feature extraction

Feature extraction runs on the ingestion server immediately after payload decryption. The extraction layer converts the raw event stream into a feature vector. The feature families are:

Keystroke dynamics features: inter-key interval distributions for the username field and password field separately (we compute the full IAT vector, mean, standard deviation, skewness, kurtosis), key hold durations, the ratio of hold to flight time, and a comparison against the account's historical keystroke signature if one exists.

Mouse and touch kinetics features: trajectory smoothness (deviation from straight-line path across all cursor segments), speed distribution, acceleration profile, the presence of ballistic movements versus guided movements, and touch pressure variance on mobile.

Session context features: dwell time on each page before the authentication action, scroll behavior on pre-login pages, time-of-day normalized against the account's historical session timing distribution, referral path, and device fingerprint consistency against account history.

Feature extraction adds approximately 3-5ms. The output is a 340-dimensional feature vector.

Stage 3: Historical feature lookup

For returning accounts, the feature vector is compared against the account's stored behavioral profile. The profile is a compressed representation of the account's historical feature distribution, updated incrementally after each confirmed-legitimate session. Profile lookup is a read from our low-latency feature store, which runs on NVMe-backed infrastructure with sub-millisecond read latency. The output is a set of comparison scores: how far the current session's features diverge from the account's historical distribution across each feature family.

For new accounts with no history, this stage returns null for the historical comparison and the model falls back to cross-account behavioral norms as the reference distribution. This is why first-session scores are generally lower for new accounts: the model has less evidence to work with.

Stage 4: Model inference

The model receives the feature vector plus the historical comparison scores and outputs a trust score. The architecture is a gradient-boosted ensemble with a calibration layer that maps the raw output to a well-calibrated probability. The model runs on a CPU inference cluster, not GPU: our feature vectors are 340-dimensional, which is small enough that batched CPU inference outperforms GPU inference at this feature size when accounting for the round-trip latency of GPU I/O.

Inference adds 8-15ms per score request. The model is updated weekly; the feature store and inference cluster are maintained in sync on a synchronized update schedule.

Stage 5: Score assembly and response

The score assembler takes the model output and the request metadata, applies any platform-specific policy overrides (if configured), and serializes the response. The response includes the trust score, a confidence band, and a set of signal explanations that describe which feature families had the most influence on the score. The serialized response is typically 180-400 bytes.

Total pipeline latency is the sum of the stages: 8-20ms transmission + 3-5ms extraction + 1ms profile lookup + 8-15ms inference + 2ms assembly = 22-43ms server-side processing, plus the network return trip of 8-20ms. End-to-end latency for the score response from SDK payload transmission to score receipt is 30-63ms, with median around 45ms. We target 100ms as our SLA ceiling at p99.

More from the blog