Columnar's Jevaro streams TypeSafe Jev answers as Apache Arrow

Columnar's Jevaro streams TypeSafe Jev answers as Apache Arrow

Columnar, the company behind Columnar Gateway, has published a blog post asking what would happen if TypeSafe AI's Jev model spoke Apache Arrow. Jev is described as a new model for turning natural language and application state into typed decisions: the developer supplies context and defines the possible answers, and Jev returns choices, scores and probabilities that code can use directly. The API delivers those answers as JSON.

The post first sets Jev in context. Other tools, such as Outlines from .txt, use constrained decoding to make existing language models produce schema-conforming output. TypeSafe took a different route. Per its announcement, as the post relays it, the company uses a new model architecture, a parallel sampler and a training method called Reinforcement Learning for Calibrated Decisions. Jev produces probabilities in parallel instead of generating an answer token by token, and TypeSafe reports substantial gains in speed and cost compared with general-purpose LLMs in its decision workflows. The post relays these claims and does not verify them. It also lists four patterns from TypeSafe's docs: speculative fan-out (many questions in one call, including speculative ones), confidence-gated routing (a different path when Jev is unsure), composite scoring (several dimensions of judgment folded into one score) and intent routing (classifying what a user wants and sending it to the right handler).

Columnar's question was whether Arrow, which it previously pitched in a post called "Stop paying the JSON tax", could speed up such pipelines. It designed an Arrow schema for Jev's three question types. Each question's answers become an Arrow column whose type is known from the question definition before inference starts. A Choice is a struct of a uint8 index into shared labels, a float64 confidence and a fixed-size list of float64 probabilities; labels sit in the field metadata. A Score is a struct of a float64 score, a float64 confidence and a fixed-size probability list, with a legend in the metadata. A Noul is a single float64 probability. Fixed-size vectors need no per-row offsets, and 64-bit floats preserve the values returned by the TypeSafe Python SDK. The authors say that if the API returned Arrow using this schema, tools such as pandas, Polars, DuckDB and Apache DataFusion could consume results without first deserializing JSON and rebuilding typed columns. They note that Databricks, Snowflake and ClickHouse can already return query results in Arrow format over HTTP, and that Hugging Face Datasets uses Arrow internally.

Jev does not offer Arrow output today, so the authors set out to simulate it. That exposed a second problem: a TypeSafe API request holds one state and a map of questions, which suits asking many questions about one live interaction (intent, urgency, refund eligibility, fraud signs) but has no native batch operation for the same questions over many independent states. Validating Jev against a sample of historical customer interactions therefore means thousands or millions of separate calls.

To work around this, Columnar built Jevaro, a small Python proxy server with Python and JavaScript clients. It accepts multiple states and a shared questions map in one request, calls Jev for each state and returns an Arrow IPC stream. The schema goes out immediately and answers follow in input order. Throughput tuning took several rounds: a long-lived HTTP/2 connection pool, a window of pending asynchronous requests refilled before result batches are written, and SDK retries that recover dropped connections and back off on rate limits and overload responses. Even so, every state still costs one upstream HTTP request and one JSON response, which Jevaro parses into Arrow arrays. The authors say they would expect a native bulk call returning an Arrow stream to achieve orders of magnitude better throughput, and at a minimum to shift the bottleneck from API overhead to inference. Choosing a proxy gives one implementation of concurrency, retries, ordering and serialization, and lets browser clients work without receiving the TypeSafe key; the cost is an extra process and network hop.

To try it, a user sets TYPESAFE_API_KEY, runs uvx jevaro-server, and calls the client. In Python, client.system_one takes a list of states and a questions map and yields a PyArrow RecordBatchReader; in JavaScript, npm install jevaro provides a client that returns an Apache Arrow AsyncRecordBatchStreamReader and also works in browsers.

For the benchmark, the authors used a batch of 10,000 synthetic customer messages, each evaluated with a Choice for department, a Score for urgency and a Noul for whether a refund was requested. The fastest run returned all 10,000 rows in 21.5 seconds, about 464 states per second. The client got the schema after 34 milliseconds and the first answer row after 290 milliseconds. The elapsed time includes upstream calls, retries and receiving the full Arrow result locally; saving to disk happens afterward. The estimated cost, based on reported token usage and TypeSafe's published pricing, was $0.20 for 30,000 answers. The authors caution that TypeSafe's rate limits are changing often, so throughput may differ.

Their conclusion is that Jev handled the workload inexpensively and relatively quickly given the overhead of 10,000 separate calls, but that this does not show how much faster a bulk Arrow API would be. They say they would like to test a native batch endpoint with the TypeSafe team. Their wish list for the TypeSafe API is narrow: a bulk endpoint that returns Arrow. Next for Jevaro are Arrow input, better adaptation to changing rate limits and experiments with output record batch sizes. The authors also say they are building novel capabilities for mixing probabilistic decisions with deterministic data processing into Columnar Gateway, and they invite readers to try Jevaro and sign up for early access to Gateway.

Key facts

  • Jevaro is an experimental Python proxy from Columnar, with Python and JavaScript clients, that takes many states and a shared questions map in one request and returns an Apache Arrow IPC stream in input order.
  • The proxy exists because the Jev API returns JSON and has no bulk endpoint; Jevaro still makes one upstream API call per state and parses each JSON response into Arrow arrays.
  • Test on 10,000 synthetic customer messages (a Choice, a Score and a Noul each): the fastest run took 21.5 seconds, about 464 states per second, with the schema arriving after 34 ms and the first row after 290 ms.
  • Estimated cost was $0.20 for 30,000 answers, based on reported token usage and TypeSafe's published pricing.
  • The authors' only ask of TypeSafe is a bulk endpoint that returns Arrow; they expect it would improve throughput by orders of magnitude, but that is an expectation, not a measurement.

Why it matters

Jev is pitched as a way to get typed decisions, probabilities and confidence from a model directly into code, and TypeSafe reports large speed and cost gains over general-purpose LLMs (a claim the post relays but does not check). If that holds, decision pipelines will move a lot of structured data, and the format it travels in starts to matter. This post is an early, practical sketch of what an analytics-friendly interface to Jev could look like: answers as typed Arrow columns that pandas, Polars, DuckDB or DataFusion can read without a JSON detour. It also names a gap that any team validating Jev on historical data would meet: no batch operation for many states.

Who it affects

Data engineers and analysts who want to run a fixed set of Jev questions over a table of historical states, for example to compare Jev's answers with known outcomes before trusting it in a live application. It also touches TypeSafe, whose API is the subject of the wish list, and readers weighing Columnar Gateway, which the authors say will include capabilities for mixing probabilistic decisions with deterministic data processing.

How to use it

Set the TYPESAFE_API_KEY environment variable and start the server with uvx jevaro-server. From Python, run uv run --with jevaro on a script that opens a TypeSafeClient pointed at http://127.0.0.1:8000, calls client.system_one with a list of states and a questions map (for example a Noul asking whether a refund is requested), and reads the PyArrow RecordBatchReader; reader.read_all() gives a table. From JavaScript, run npm install jevaro and use the client's systemOne method, which returns an Apache Arrow AsyncRecordBatchStreamReader and also works in browsers. Each question type maps to an Arrow column: Choice and Score are structs with confidence and a fixed-size probability list, and Noul is a plain float64. The source gives no price for Jevaro itself; the cost figure covers Jev usage.

How solid is it

The numbers come from the authors' own test, and they are clear about its limits. The 21.5 seconds is the fastest run, not an average, and the post gives no number of runs and no hardware. The $0.20 is an estimate from reported token usage and published pricing, not an invoice. The benchmark compares no baseline: there is no measured comparison with a native bulk endpoint, with plain JSON handling or with other models, and no accuracy or calibration results for Jev are reported. The claim that a native Arrow bulk call would bring orders of magnitude better throughput is an expectation. The post comes from a company that sells a related product.

Risks and caveats

Jevaro is a simulation of an interface that does not exist: the Jev API does not offer Arrow output today, and TypeSafe has not announced or agreed to a bulk or Arrow endpoint; the authors say only that they would like to test one with the TypeSafe team. Every state still triggers a separate upstream request and JSON response, so the conversion overhead the authors want to remove is still there. TypeSafe's rate limits are changing often, so throughput may differ from the figures reported. Running the proxy adds a process and a network hop, though the authors note the same optimizations could run in a client. The post also relays TypeSafe's speed and cost claims without verifying them.

“our wish list for the TypeSafe API is narrow: a bulk endpoint that returns Arrow.”

— Columnar, "What if Jev spoke Arrow?"