Avaloka 1.0 · Open source · Apache 2.0

Introducing Avaloka

A virtual team of AI-enabled data scientists that turns raw data and an open question into analysis you can inspect, challenge, and run again.

Any Cloud. Any Model. Better Data Science.
September 21, 2026 · Avaloka team · 9 minute read
How Avaloka works● Avaloka 1.0
InputA question + connected dataFiles, databases, object stores, documents, and APIs
TeamAI data science team
PlanProfileCodeValidateTrainVerify
OutputVerified decision evidenceResults, baselines, caveats, code, artifacts, and provenance
Any CloudAny ModelHuman-led
Crisp architecture view

One question. One coordinated system.

Avaloka connects enterprise context to a team of specialized agents, routes the workload to the right compute, and returns evidence that can be reviewed.

Workload routeLocalOn-premAny cloudRay + Kubernetes
Why we built Avaloka

The data was never the whole problem.

In 2007, database pioneer and Turing Award laureate Jim Gray described a fourth paradigm of science: discovery shaped by the ability to explore vast, messy data.

As data scientists, we share that ambition. We also know what the data deluge revealed. Enterprises can move petabytes through distributed systems and still struggle to answer a precise business question with evidence that a statistician would defend.

Galileo using a telescope, Newton studying orbital models, and Babbage working with a mechanical analytical engine
ObservationGalileo turned a telescope toward the sky.ExplanationNewton gave motion a mathematical model.ComputationBabbage imagined machinery that could process instructions and data.

The limiting factor is rarely one more cluster. It is the work around the computation: finding the right data, understanding its grain, connecting facts across systems, choosing a valid method, checking assumptions, interpreting uncertainty, and turning the result into something a decision maker can use. Ten teams working in tandem can still produce ten partial views of the same enterprise.

The hard part is not processing more data. It is finding the signal that matters for this decision, then showing why it should be trusted.
01

Processed is not understood

A warehouse can make data available. It cannot decide whether a join changes the grain, whether a metric answers the question, or whether an apparent pattern will survive a serious statistical test.

02

The enterprise picture is fragmented

Customer truth lives across transactions, product events, support records, documents, models, and people. Code built against one visible table can be correct and still miss the business.

03

The scientist remains the orchestrator

Notebooks and chat interfaces help, but repeatedly approving or editing lines of code still leaves the human coordinating data engineering, analysis, validation, infrastructure, and reporting.

Jim Gray’s vision is our starting point

Gray wanted better tools so scientists could spend more time on discovery and less time on data preparation. Avaloka applies that idea to enterprise data science. Read the Avaloka preface and the reference on The Fourth Paradigm.

From code assistance to investigation ownership

A model with a chatbot is not sufficient to solve data engineering and data science problems.

Avaloka is an open-source Data Science Team of AI agents.

A frontier model can write a useful function, explain a chart, and suggest an analytical path. JupyterHub is a valuable workspace. Code assistants are valuable copilots. But a real investigation does not end when a code cell looks plausible.

It must assemble the relevant context, decide what to compute, generate production-scale work, validate it against the real schema and execution environment, run it at the right scale, test the result, and preserve the evidence. That is a team problem.

A coding chatbot

  • Works with the context currently visible
  • Suggests code and asks for approval
  • Optimizes the next response
  • Often stops at prose, a cell, or a script

Avaloka’s AI team

  • Builds a connected view of the question
  • Plans, writes, validates, and executes
  • Checks data, models, and final claims
  • Leaves code, artifacts, evidence, and provenance

In Avaloka 1.0, a Planner coordinates specialists for sampling, profiling, coding, validation, execution, data transfer, model training, visualization, and claim verification. In Avaloka Enterprise, a knowledge graph connects operational data, documents, prior work, and derived artifacts so the analysis can move toward a true 360 degree view instead of treating every prompt as an isolated coding task.

Avaloka 1.0

Begin with the question. Build the stack beneath it.

Avaloka works from intent to evidence. The human frames the goal and reviews consequential choices. The AI team coordinates the analytical work underneath.

01 / ConnectQuestion plus enterprise context

Files, databases, object stores, APIs, and the business meaning needed to interpret them.

02 / InvestigateSpecialists share one mission

Profile, sample, plan, code, validate, execute, visualize, train, and evaluate.

03 / ReviewDecision evidence

Results with code, checks, baselines, caveats, artifacts, and provenance.

CONNECT

Work across data boundaries

Avaloka can query files and databases where they live. Its Data Transfer Agent can move and validate data across databases, accounts, and clouds, with commercial connectors for AWS, Azure, and Google Cloud.

SCALE

Use the smallest faithful workload

For exploration, Avaloka samples for fast approximate answers and labels the result as sampled. For larger work, it can use database statistics, optimizer plans, physical partition layouts, Ray, and Kubernetes. Avaloka Enterprise can schedule approved jobs on customer-controlled cloud infrastructure and notify users when they complete.

CODE

Generate production-scale code, then challenge it

The coding agent writes pseudocode before executable code. Verifiers apply three layers of validation: syntax, schema-aware functional checks, and logical verifiability, including optimizer-aware checks when the data source exposes them.

SERVE

Train, compare, and serve models

Commercial capabilities can plan and run training, screen for leakage, compare against a simple baseline, register artifacts, and provide prediction services on Kubernetes.

Evidence before eloquence

Trust is earned before the answer appears.

A fluent explanation is not proof. Avaloka separates the checks because data integrity, model quality, code safety, and claim support fail in different ways.

DATA

Integrity

Check grain, schema, missingness, target leakage, duplicate rows across splits, identifier-like features, and time-related leakage before trusting an analysis or training run.

CODE

Execution safety

Validate syntax, schema compatibility, likely fanout, dangerous full scans, and bounded sample behavior before a production-scale execution.

MODEL

Evaluation

Choose a split strategy that fits the data, record why it was chosen, and compare every trained model with a mandatory simple baseline.

CLAIM

Verification

Check the final narrative against computed evidence for invented numbers, contradictions, causal overreach, and unsupported certainty.

Open source is inspectable, but inspectability is not the same as a proper benchmark. Avaloka publishes the checks and test evidence available today. The open-source system has not yet been comprehensively benchmarked, and measuring it openly is a near-term goal.

The public technical guide documents an internal conformance suite of 281 prompts across 13 capability areas. That is useful coverage evidence, but it is not an end-to-end performance benchmark.

Vision

From an answer to a Decision Packet. Avaloka already preserves code, artifacts, metrics, and provenance. The future Decision Packet brings the question, evidence, caveats, approvals, and action record into one reviewable object.

Research benchmark

Measure the work that makes data science reliable.

The Avaloka research protocol tests robustness and ML productivity, not just whether a model can produce a plausible query.

Failures preventedunsafe executions stopped per 100 tasks
Plan regression detectioncost and cardinality changes caught before execution
Schema drift sensitivityrenames, type changes, nullability, and enum expansion
Change safetybackward-compatible plans with rollback
Cost containmentfull-data executions avoided through sampling and preflight checks
Time to productiontime to trainable data and time to an inference endpoint
Published evaluation protocol
Task suitePrimary metricsComparison
ExplorationPrevented execution failures and cost containmentSingle-shot natural-language-to-SQL, agent validation without optimizer signals, and Avaloka with optimizer-aware validation plus sampling
OperationsSchema-drift sensitivity and change-safety scoreThe same approaches under controlled schema and migration changes
SREPlan-regression detection and full-data executions avoidedReactive execution versus preflight estimation and bounded trials
ML productivityTime to trainable data, manual steps eliminated, and time to endpointManual feature pipeline versus Avaloka’s profiling-led path

The public research paper defines this protocol but does not report completed scores. We will publish measured results rather than present planned metrics as benchmark outcomes.

Avaloka Enterprise engagements also provide deployment-specific benchmarks that compare cycle time, answer quality, governance trace, and cost with the work required from world-class data science teams building and serving production models.

Open by default

Cloud-agnostic by design.

Run Avaloka on a laptop, on infrastructure you control, or on Kubernetes and Ray in the cloud you choose. The open-source system connects to clusters your configuration can reach. Avaloka Enterprise adds managed provisioning, commercial data connectors, scheduling, notifications, training, and inference services.

Local + on-premPrivate infrastructure
AWSCustomer cloud
Google CloudCustomer cloud
Microsoft AzureCustomer cloud

Model choice is configuration, not lock-in. Use hosted providers, cloud model services, OpenAI-compatible endpoints, or local models, and change that choice as privacy, latency, cost, and capability requirements change.

The Avaloka promiseAny Cloud. Any Model. Better Data Science.
Inspect the system for yourself.

The source, architecture, guides, supported formats, tests, and edition boundaries are maintained with the code.

Open GitHub ↗
Future vision

Make evidence continuous.

The next chapter is not simply more agents. It is a tighter relationship between the question, the experiment, the evidence, and the decision.

01

Swarm Intelligence

Specialists compare approaches, share evidence, surface disagreement, and converge on a stronger analytical path.

02

Workload costing and optimization

Route work across samples, databases, local compute, and distributed clusters by cost, latency, fidelity, and confidence.

03

Ability to Experiment

Move from observation to explicit hypotheses, treatments, controls, measured outcomes, and reviewable decisions.

04

New domains

Add domain specialists and methods without rebuilding the data science system around each new problem.

05

Multimodal dataset analysis

Investigate tables together with documents, images, audio, and video while keeping the evidence connected.

06

Physical AI dataset understanding

Reason across sensor data, telemetry, simulation, and embodied-system behavior with the same verification discipline.

These are future directions, not Avaloka 1.0 feature promises.

Build it with us

Better data science begins with verifiability and evidence.

Try Avaloka on data you understand. Ask a difficult question. Read the plan and the generated code. Challenge the statistics. Inspect the evidence. Help us build the system Jim Gray’s vision deserves.

Any Cloud · Any Model · Better Data Science

Start with a question. Arrive at an insight, make a decision.

Explore Avaloka on GitHub ↗