VizSoup

About VizSoup and its methods

VizSoup is a set of small, free data tools that run entirely in your browser. This page says how each one works, which conventions it follows, and what it does not do.

How the tools work

Every tool is JavaScript that runs on your device. There is no account, no server-side processing and no upload: pasted text and opened files are read in the browser tab and forgotten when you close it. Charts are built as SVG text by small functions written for this site, with no third-party chart library, and the same functions draw the example you see before you type anything. The site sets no cookies of its own; the privacy page describes exactly what this build includes.

Reading your data

All the CSV tools share one parser that follows RFC 4180: fields in double quotes may contain the delimiter, line breaks and doubled quotes. A UTF-8 byte-order mark is ignored. The delimiter (comma, semicolon, tab or pipe) is detected by trying each and keeping the one that gives the most consistent column count without malformed quoting; you can override it. Whether the first row is a header is detected too: a first row made entirely of numbers is treated as data.

Numbers are read as people type them in spreadsheets: thousands separators in groups of three, a leading currency sign ($, €, £, ¥), a trailing percent sign (read as the number itself, so 12% is 12), and parentheses for negatives. Anything else, including blank cells, is not a number and is skipped rather than counted as zero; each tool tells you how many cells it skipped. Decimal commas, as in 3,5, are not converted by the chart and statistics tools; run the data through the CSV cleaner first, which converts them.

Cleaning

The CSV cleaner applies a list of steps in order to a copy of your table, starting again from your input after every change, so switching a step off undoes it exactly. Each step reports how many cells, rows or columns it changed, and the preview highlights every changed cell against the original. It runs in a Web Worker, a background thread of the page, so tables of 50,000 rows and more do not freeze the tab.

Two steps have to interpret values, and both refuse to guess. The date step decides whether a column is day-first or month-first from dates that can only be read one way (15/04/2024 can only be 15 April); if a column has none, it asks. Dates are checked against the calendar, dotted dates (21.03.2024) are read day-first, and two-digit years follow the POSIX rule: 69–99 is the 1900s, 00–68 the 2000s. The number step decides between a decimal point and a decimal comma the same way, from values such as 3,5 or 1,234.50 that allow only one reading, and asks about a column of values like 1,234. It rewrites digits as text rather than through floating point, so no precision is lost and leading zeros are kept.

Charts

Bar and line charts, the pie and donut maker, the histogram maker and the scatter plot share these rules:

Statistics

The summary statistics calculator uses the standard definitions: the sample standard deviation divides by n − 1 and the population version by n; skewness is the adjusted Fisher–Pearson coefficient that Excel's SKEW returns. Sums use compensated (Kahan) summation to avoid floating-point drift.

Quartiles are shown under three conventions because software disagrees. Hyndman and Fan (1996) catalogued nine definitions in use:

Outliers use Tukey's fences, 1.5 × IQR beyond the first and third quartiles, with the quartile convention of your choice. The fences flag values worth checking; they are not a rule for deleting data.

The histogram maker computes five bin rules side by side: Sturges (1926), Freedman–Diaconis (1981), Scott (1979), square root and Rice. Bins include their lower edge and exclude their upper edge, except the last. By default the width is rounded to a readable number, which can change the bin count by one or two.

The scatter plot fits an ordinary least-squares line of y on x and reports the Pearson correlation r, r², the residual standard error and the slope's standard error and t statistic. The verbal labels for r (negligible, weak, moderate, strong at 0.1, 0.3 and 0.5) follow Cohen's conventions and are a rough guide only. The sample data there is Anscombe's quartet dataset I (F. J. Anscombe, The American Statistician, 1973).

Summaries and conversions

The pivot table groups rows by exact label (after trimming spaces) and computes sum, count, mean, median, minimum, maximum or distinct count. Totals always come from the underlying rows, never from the summary cells, so a total average is a true overall average.

The table converter writes GitHub Flavored Markdown pipe tables, escaping pipes and backslashes, and semantic HTML with every value entity-escaped. The JSON and CSV converter flattens nested objects into dotted column names and only converts a CSV value to a JSON number when that is lossless, so ZIP codes with leading zeros and IDs longer than about 15 digits stay as strings.

Sample data

Each tool opens with example data so the page shows a working result. Apart from Anscombe's quartet, the samples (shop orders, a household budget, delivery times, a sales ledger, product and user records) are invented for illustration, and each page says so. None of them is a real measurement.

What the tools do not do

Apart from the cleaner's date step, they do not parse dates. They do not fit curves, run hypothesis tests beyond the slope's t statistic, or remember anything between visits. They are designed for tables up to tens of thousands of rows; beyond that, a spreadsheet or a statistics package is the better tool.

Who makes this

VizSoup is a small independent site, built so that every result can be checked: each page names its method and its limits. It carries no sponsored rankings. If it shows ads or links that earn a commission, they are labelled where they appear, and they have no influence on how any tool works.

Corrections

The methods are stated so they can be checked. If a result here disagrees with another tool, the most likely cause is a different convention (quartiles, sample versus population standard deviation, or bin edges), and the relevant page explains which one is used here.