Open-source data analysis workspace

Data analysis you can check.

Upload a spreadsheet, survey export or research dataset. Explain Your Data profiles it, helps you clean it, runs the right statistics, draws the charts and explains the result in plain language. AI decides what to run and writes the explanation. Code computes every number, and a validator checks the answer before you see it.

The live app runs on free hosting that sleeps when idle, so the first visit can take up to a minute to load. Three sample datasets are built in if you don't have a file to hand.

explain-your-data.onrender.com · retail_sales_2025-26.csv
The Overview screen for a retail sales dataset: total revenue $1.39M, 3,532 orders, a dataset health score of 94 out of 100, ranked findings and suggested data fixes.
CSVExcelODSJSONZIPDOCXPDFTXT · MD · HTML · XMLGoogle Forms exportsSPSS · Stata · SAS · Parquet (server)
The principle

The model never does the arithmetic.

Language models are good at choosing an analysis and explaining it. They are unreliable calculators. So the work is split, and each half does what it is good at.

01AI

Plan

Reads your question and a summary of the columns (never the rows) and picks tools: aggregate, regression, forecast, chart.

02code

Compute

Deterministic code runs those tools on the full dataset in your browser: descriptive statistics, tests, models, forecasts.

03code

Verify

Every number in the drafted answer is matched against the computed results. Unsupported figures are sent back once for repair, then flagged.

04AI

Explain

The checked result is written up in plain language, or rewritten for a CEO, a client, a student or an academic paper.

Without an AI key the app still works end to end: a built-in engine maps questions to the same analyses, so every feature except AI-written prose runs offline in the browser.

The validator

A number that wasn't computed doesn't get through silently.

Extracts every figurePercentages, currency, thousands separators and abbreviations like $1.24M or 54.3K.
Matches against the tool resultsAllows for rounding (about 1%), 18.4% stored as 0.184, and a sign dropped in prose ("falls 0.28" for a coefficient of −0.28). Years and small counting words are ignored.
One repair round, then a visible flagThe model gets one chance to fix the answer. Anything still unsupported is shown to the user as "not verified" rather than hidden.
Same rules on server and browserA Python implementation guards the API; a JavaScript copy checks answers produced inside the browser. The two give identical results on the same test cases.
Question
Which region has the slowest delivery, and how much slower is it?
Computed by code
aggregate(mean delivery_days by region)
Perth      4.28
Melbourne  2.50
Sydney     2.35
…
Adelaide   2.08
Drafted answer
Perth is slowest at 4.28 days on average, against 2.08 in Adelaide: about 106% longer.
Not verified: 106%. This figure was not found in the computed results, so treat it with caution.
Honest by default

It tells you when a result shouldn't be trusted.

Most dashboards will happily report a 1,780% jump caused by one typo. These checks run on every upload, before any chart is drawn.

One row dominating

If a single value is more than a fifth of a total, headline totals, trends and segment shares are marked low-confidence and the row is surfaced first.

"One row makes up 58% of all revenue"

Impossible values

Negative prices, quantities, ages or durations are caught, counted in the health score and blanked (not deleted) with one click.

"1 impossible negative value in delivery_days"

Outliers in skewed data

Money and counts are skewed, so a plain IQR rule flags hundreds of normal sales. Skewed columns are also checked on a log scale to isolate the real anomaly.

"1 extreme outlier in revenue (max 48,250 vs median 45.9)"

Trends that aren't trends

A higher second half is not called growth unless the linear trend is statistically significant. The wording says so.

"higher in the second half (24.5%), but no clear trend"

Forecasts with real uncertainty

Prediction intervals are never narrower than the model's own error on held-out months, and never go below zero for quantities that can't.

"Jan 2026: $4,920 (80% interval $2,579–$7,262)"

Segments that aren't segments

Clustering rejects solutions where a group is a handful of unusual rows, and says "no clear segments" instead of inventing personas.

k-means on winsorised features, minimum group size
Tour

Upload, clean, analyse, ask, report.

Discover

The Overview screen on a phone.
The Discover screen on a phone.

Works on a phone

Every screen reflows to a single column, the workspace navigation becomes a scrolling tab bar, and the question box stays within thumb reach.

What's inside

Built for research data, surveys and business data.

Profile and clean

  • Type detection for numbers, dates, currencies (incl. € and decimal commas), IDs and Likert scales
  • 0–100 health score with explained issues
  • 14 replayable cleaning steps with preview, undo and lineage

Statistics

  • Welch and paired t-tests, ANOVA, chi-square, Mann–Whitney, Kruskal–Wallis
  • OLS regression with coefficients, p-values and R²; correlations with confidence intervals
  • Forecasting with holdout model selection, k-means, PCA, Cronbach's α

Charts

  • 18 chart types from a builder or a sentence ("monthly revenue by region")
  • A chart validator that flags poor choices and suggests a fix
  • Dashboards with shared filters; PNG and SVG export

Ask

  • Questions answered with the same verified analyses
  • "Why", "what changed in December" and ranking questions routed to drivers, change decomposition and averages
  • Follow-ups and "explain it like…" rewrites

Reports

  • Five styles: executive, academic, technical, student, marketing
  • Export to HTML, PDF and Markdown
  • Every chart carries the calculation that produced it

Reproducible

  • Clean data to CSV, Excel or JSON
  • A pandas script and an R script that replay every cleaning step on the original file
  • The Python script reproduces the app's cleaned output exactly on all 11 files in our test set
Architecture

Rows stay in the browser.

The analysis engine runs client-side. The server only ever sees a column summary and computed results, which is what the AI needs to plan and explain.

Browser
vanilla JS · Plotly
  • File readers, type inference
  • Cleaning, statistics, charts
  • Validator (JS copy)
Next.js
apps/web
  • Serves the workspace
  • Proxies AI and file conversion
  • Keys stay server-side; CSP headers
FastAPI
services/api · pandas · SciPy
  • AI router (Claude / GPT) with fallback
  • Number verifier and repair round
  • SPSS, Stata, SAS, Parquet conversion

File contents are treated as untrusted data in every prompt, so instructions hidden inside a spreadsheet cell are not followed.

Status

What works today, and what's next.

Live
  • Upload of 20+ formats, including ZIP bundles and Google Forms exports
  • Profiling, health score, cleaning with lineage, Python/R export
  • 15 analyses, 18 chart types, dashboards, five report styles
  • Ask with validated answers; built-in engine when no AI key is set
  • Mobile layout, dark mode; API tests and an evaluation harness
Next
  • Saved projects and version history (Postgres schema is in the repo)
  • Sign-in, public share links and embeddable charts
  • DOCX and PPTX report export
  • Google Sheets, Drive and database connectors
  • Larger-than-memory files processed on the server

Try it with your own data, or one of the samples.

No sign-up. Files are read in your browser.