Scatter plot with linear regression
Paste two columns of numbers. You get the scatter plot, the least-squares line through it, and the statistics that describe how well the line fits, each explained below.
Least-squares line
y = 3.000 + 0.5001x
r = 0.816 (strong positive), r² = 0.667: the line accounts for 66.7% of the variation in y. At x = 10 the line predicts y = 8.001.
Points used
11
Slope
0.5001
Intercept
3.0001
Pearson r
0.8164
r²
0.6665
Slope std. error
0.1179
t (slope = 0)
4.241
Residual std. error
1.2366
How to use it
Paste a table with at least two numeric columns and choose which is x and which is y. Put the variable you want to explain or predict on the y axis. Rows where either value is blank or not a number are skipped, and the page says how many. Enter an x value to read the line's prediction at that point.
The method: ordinary least squares
The line is the one that makes the sum of squared vertical distances from the points to the line as small as possible. With x̄ and ȳ the means of each column:
slope b = Σ(x − x̄)(y − ȳ) / Σ(x − x̄)²intercept a = ȳ − b·x̄r = Σ(x − x̄)(y − ȳ) / √(Σ(x − x̄)² · Σ(y − ȳ)²)r² = 1 − SSE / SST, the share of the variance in y explained by the lineresidual standard error = √(SSE / (n − 2)), the typical vertical miss, in y's unitsSE(b) = residual SE / √Σ(x − x̄)²andt = b / SE(b)
SSE is the sum of squared residuals, and SST is the total sum of squares of y around its mean. For a straight-line fit with an intercept, r² computed this way equals the square of r.
Worked example: Anscombe's quartet
The sample is the first of four small datasets the statistician Francis Anscombe published in 1973. Its 11 points give a mean x of 9 and a mean y of 7.50. The fitted line is y = 3.00 + 0.500x, with r = 0.816 and r² = 0.667, so the line accounts for two thirds of the variation in y. The slope's standard error is 0.118, giving t = 4.24 on 9 degrees of freedom. At x = 10, the line predicts 8.00.
Anscombe's point was that the other three datasets in the quartet give almost exactly the same line, the same r and the same r², while looking completely different when plotted. One is a smooth curve, one is a perfect line with one outlier, and one has every point but one at the same x. The statistics alone cannot tell them apart. That is the reason this tool always draws the points and never reports a fit without them.
Reading r
| |r| | Label used here | r² |
|---|---|---|
| below 0.1 | negligible | < 1% |
| 0.1 to 0.3 | weak | 1–9% |
| 0.3 to 0.5 | moderate | 9–25% |
| 0.5 and above | strong | 25% + |
These cut-offs follow Jacob Cohen's widely used conventions for the behavioural sciences. They are a starting point for describing a result, not a verdict: what counts as strong depends on the field and on how noisy the measurements are.
Where it stops being reliable
- Curves. A straight line fitted to a curved pattern can still give a high r². Look at the plot: if the points bend away from the line at both ends, the line is the wrong model.
- Outliers. Least squares squares the distances, so a single far-off point can pull the line and inflate or destroy r. Try the fit with and without it.
- Extrapolation. The prediction is only supported inside the range of x you have. Beyond it, the line is an assumption.
- Causation. Correlation between two columns can come from a third variable that drives both. Nothing in the arithmetic can detect that.
For the distribution of either column on its own, use the summary statistics calculator or the histogram maker.
Common questions
What is a good r² value?
It depends entirely on the field. In a controlled physics measurement, r² below 0.99 may suggest a problem; in social science, 0.2 can be a meaningful finding. r² is the share of the variation in y that the straight line accounts for. It says nothing about whether the relationship is causal, or whether a straight line is the right shape.
What is the difference between r and r²?
r, the Pearson correlation coefficient, runs from −1 to 1 and its sign gives the direction of the relationship. r² is its square, from 0 to 1, and has a direct reading: the proportion of the variance in y explained by the line. An r of 0.5 sounds substantial but explains only a quarter of the variance.
Does it matter which column is x and which is y?
For r and r², no: they are symmetric. For the line, yes. Least squares minimises vertical distances, so regressing y on x and x on y give two different lines unless the points are perfectly aligned. Put the variable you want to predict on the y axis.
Can I fit a curve instead of a straight line?
Not with this tool. A common workaround is to transform a column before pasting it: taking the logarithm of y turns exponential growth into a straight line, and logging both columns fits a power law. The slope then has a different meaning, so interpret it with care.
What does the t value next to the slope mean?
It is the slope divided by its standard error, the usual test statistic for whether the slope differs from zero, with n − 2 degrees of freedom. As a rough guide, with more than about 30 points an absolute t above 2 corresponds to the conventional 5% significance level. It assumes independent points with roughly normal, constant-spread residuals.
Other tools
- CSV cleaner How do I clean up this messy CSV export?
- CSV to bar or line chart How do I turn this spreadsheet into a chart quickly?
- Pie and donut chart maker What share of the total is each category?
- Histogram maker How many bins should my histogram have?
- Summary statistics calculator What are the mean, median and outliers of this data?
- Pivot table and group-by summarizer What are total sales by region and product?
- CSV to Markdown and HTML table converter How do I put this spreadsheet in a README or web page?
- JSON to CSV converter (and back) How do I open this JSON export in a spreadsheet?