VizSoup

Scatter plot with linear regression

Paste two columns of numbers. You get the scatter plot, the least-squares line through it, and the statistics that describe how well the line fits, each explained below.

Sample: dataset I of Anscombe's quartet (1973), a published teaching dataset.

Least-squares line

y = 3.000 + 0.5001x

r = 0.816 (strong positive), r² = 0.667: the line accounts for 66.7% of the variation in y. At x = 10 the line predicts y = 8.001.

Points used

11

Slope

0.5001

Intercept

3.0001

Pearson r

0.8164

r²

0.6665

Slope std. error

0.1179

t (slope = 0)

4.241

Residual std. error

1.2366

y against xScatter plot of y against x, 11 points.4681012y468101214x(10, 8.04)(8, 6.95)(13, 7.58)(9, 8.81)(11, 8.33)(14, 9.96)(6, 7.24)(4, 4.26)(12, 10.84)(7, 4.82)(5, 5.68)

How to use it

Paste a table with at least two numeric columns and choose which is x and which is y. Put the variable you want to explain or predict on the y axis. Rows where either value is blank or not a number are skipped, and the page says how many. Enter an x value to read the line's prediction at that point.

The method: ordinary least squares

The line is the one that makes the sum of squared vertical distances from the points to the line as small as possible. With x̄ and ȳ the means of each column:

SSE is the sum of squared residuals, and SST is the total sum of squares of y around its mean. For a straight-line fit with an intercept, r² computed this way equals the square of r.

Worked example: Anscombe's quartet

The sample is the first of four small datasets the statistician Francis Anscombe published in 1973. Its 11 points give a mean x of 9 and a mean y of 7.50. The fitted line is y = 3.00 + 0.500x, with r = 0.816 and r² = 0.667, so the line accounts for two thirds of the variation in y. The slope's standard error is 0.118, giving t = 4.24 on 9 degrees of freedom. At x = 10, the line predicts 8.00.

Anscombe's point was that the other three datasets in the quartet give almost exactly the same line, the same r and the same r², while looking completely different when plotted. One is a smooth curve, one is a perfect line with one outlier, and one has every point but one at the same x. The statistics alone cannot tell them apart. That is the reason this tool always draws the points and never reports a fit without them.

Reading r

|r|Label used herer²
below 0.1negligible< 1%
0.1 to 0.3weak1–9%
0.3 to 0.5moderate9–25%
0.5 and abovestrong25% +

These cut-offs follow Jacob Cohen's widely used conventions for the behavioural sciences. They are a starting point for describing a result, not a verdict: what counts as strong depends on the field and on how noisy the measurements are.

Where it stops being reliable

For the distribution of either column on its own, use the summary statistics calculator or the histogram maker.

Common questions

What is a good r² value?

It depends entirely on the field. In a controlled physics measurement, r² below 0.99 may suggest a problem; in social science, 0.2 can be a meaningful finding. r² is the share of the variation in y that the straight line accounts for. It says nothing about whether the relationship is causal, or whether a straight line is the right shape.

What is the difference between r and r²?

r, the Pearson correlation coefficient, runs from −1 to 1 and its sign gives the direction of the relationship. r² is its square, from 0 to 1, and has a direct reading: the proportion of the variance in y explained by the line. An r of 0.5 sounds substantial but explains only a quarter of the variance.

Does it matter which column is x and which is y?

For r and r², no: they are symmetric. For the line, yes. Least squares minimises vertical distances, so regressing y on x and x on y give two different lines unless the points are perfectly aligned. Put the variable you want to predict on the y axis.

Can I fit a curve instead of a straight line?

Not with this tool. A common workaround is to transform a column before pasting it: taking the logarithm of y turns exponential growth into a straight line, and logging both columns fits a power law. The slope then has a different meaning, so interpret it with care.

What does the t value next to the slope mean?

It is the slope divided by its standard error, the usual test statistic for whether the slope differs from zero, with n − 2 degrees of freedom. As a rough guide, with more than about 30 points an absolute t above 2 corresponds to the conventional 5% significance level. It assumes independent points with roughly normal, constant-spread residuals.

Other tools