// HACKER NEWS — CYBERSECURITY
What if Jev spoke Arrow?
Jev is TypeSafe AI’s new model for turning natural language and application state into typed decisions. You supply the context and define the possible answers. Jev returns choices, scores, and probabilities that your code can use directly. The API delivers those answers as JSON. If you’ve managed to avoid hearing about Jev lately, the rock you’re hiding under has excellent soundproofing.
Jev is part of a wider effort to make AI outputs easier to use in code. Other tools, such as Outlines from .txt, use constrained decoding to make existing language models produce outputs that conform to a schema. TypeSafe took a different approach. In its announcement, the company describes a new model architecture, a parallel sampler, and a training method called Reinforcement Learning for Calibrated Decisions. Jev produces probabilities in parallel, avoiding the work of generating an answer token by token. TypeSafe reports substantial gains in speed and cost compared with general-purpose LLMs in its decision workflows.
Jev is a new primitive, and nobody yet knows the full scope of what it will make possible. This post reflects our thinking at this point in time, and we expect it to evolve. But some powerful patterns are already clear. TypeSafe’s docs describe several. Speculative fan-out asks many questions in one call, including speculative ones, and lets your code decide which answers are relevant. Confidence-gated routing treats confidence as a second decision axis, so your code can take a different path when Jev is unsure. Composite scoring combines several dimensions of judgment into a single score. Intent routing classifies what a user wants and sends the request to the right handler.
Together, these patterns open the door to fundamentally probabilistic workflows and pipelines, in places where until recently we would have assumed only deterministic ones were practical. The sophistication you can achieve is astounding.
Pipelines like these move a lot of structured data, which got us curious about how Jev might work with another technology that combines structure with performance and efficiency: Apache Arrow. In Stop paying the JSON tax, we described how Arrow can speed up data pipelines by avoiding conversions to JSON and back. Could it do that here?
To explore that question, we first designed an Arrow schema to represent Jev’s answers. Jev has three question types, each with a different answer shape:
We chose to represent each question’s answers as an Arrow column. The question’s definition tells us the column’s type before inference starts. Choice labels and Score legends describe the possible answers, so we can put them in the schema’s field metadata. The predictions and probabilities go in the data buffers. Arrow extension types let us attach that semantic meaning to ordinary Arrow storage types:
N is the number of choices or score levels. Choice stores a one-byte index into the shared labels. Fixed-size probability vectors need no per-row offsets, and 64-bit floats preserve the values returned by the TypeSafe Python SDK. The schema metadata supplies the meaning of our choice indices and score levels. Together, the values and metadata are enough to reconstruct each original answer object.
The TypeSafe API returns these answers as JSON today. If it returned Arrow directly using this schema, tools such as pandas, Polars, DuckDB, and Apache DataFusion could consume the results without first deserializing JSON and rebuilding typed columns. Compatible consumers can use the Arrow buffers without copying them.
This is already a common way to exchange structured data. Databricks, Snowflake, and ClickHouse can return query results in Arrow format over HTTP. Hugging Face Datasets uses Arrow internally and lets you retrieve Arrow tables directly. Giving Jev an Arrow output option would let its answers join those same data workflows.