This is a discussion of the F-test from a regression perspective and not from a analysis of variance perspective, even though there is also an ANOVA table listing out the sums of squares resulting from linear regression. From a regression perspective, when we think of F-test, we usually think of the F-test that is to test the significance of the entire regression model, which is an “all or nothing” test. If the F-statistic is sufficiently large, then we know the entire linear regression model is useful in explaining the variability in the response, i.e., some of the predictors are significant and have an impact on the response. Otherwise, we conclude that the predictors as a group do not improve the prediction of the response over that obtained by regressing the response on the intercept coefficient (this is the intercept-only model). For convenience, we call the “all or nothing” F-test as the whole model F-test. What if we are interested in knowing whether a subset of the predictors is significant? In other words, what if the comparison is not between the full model and the intercept-only model but is between the full model and a smaller model that is larger than the intercept-only model? We propose that the whole model F-test can be tweaked to test the significance of a reduced model versus the full model.
See here for an application of the generalized F-test.
The Whole Model F-Test
First, we start with the more familiar concept of the whole model F-test. As mentioned above, this test is “all or nothing.” The “all” is the full model with all the predictors and the “nothing” is the intercept-only model. In the intercept-only linear regression model, there are no slope coefficients and there is only the intercept coefficient. With the intercept-only model, the model fitting is easy since the fitted value is simply the sample mean of the observed response values. With “all or nothing,” it is about choosing between these two propositions: either we are doing no better than using the sample mean response as the prediction (going with the intercept-only model) or the linear regression model with all predictors is useful for explaining the variation in the response (going with the full model). These two models are described as follows:
-
(1)……….
(1)……….…………………………………………(Intercept-Only Model)
The hypotheses to be tested are the following:
-
(2)……….
(2)………. for at least one
If the null hypothesis holds, then we opt for the the intercept-only model, i.e., none of the predictors
makes any difference. If there is sufficient evidence to reject the null hypothesis, then we are in favor of the full model. As we will see below, the null hypothesis corresponds to the smaller model. The evidence for or against the null hypothesis is summarized by the following F-statistic, a quantity found in the ANOVA table,
-
(3)……….
where is the sample size. See here for a more in-depth discussion of the ANOVA table and the F-statistic. Under the null hypothesis
, the F-statistic follows an F distribution with degrees of freedom
and
with expected value close to 1. An observed value of the F-statistic that is substantially larger than 1 provides evidence to support the rejection of the null hypothesis.
In (3), RSS is the residual sum of squares, which is the sum of the squares of the differences “observed response minus fitted response.” On the other hand, TSS is the total sum of squares, which is the sum of squares of the differences “observed response minus sample mean response.” RSS is the amount of variability that is left unexplained after performing the regression, while TSS is the amount of variability inherent in the response before the regression is performed. As a result, TSS – RSS is the amount of variability in the response that is explained by performing the linear regression.
RSS and TSS can be regarded as as measurements of error because the difference in each sum of squares can be regarded as an error. The difference “observed response minus fitted response” measures the error associated with a fitted value in the full model. On the other hand, the difference “observed response minus sample mean response” measures the error associated with a fitted value in the intercept-only model. Thus RSS is a sum of squares that measures the error with respect to the full model (the smaller the RSS, the better the full model fits the data). In the terminology used here, both RSS and TSS are called errors sums of squares with respect to the appropriate models (for RSS, it’s the full model and for TSS, it’s the intercept-only model).
The model with more predictors will always be able to fit the data at least as well as the model with fewer predictors. Thus, the full model will give a better fit to the data than the intercept-only model, i.e., RSS TSS. The hypothesis test described in (2) is to determine whether the full model will provide a significantly better fit to the data than the intercept-only model, i.e., whether RSS will be significantly lower than TSS. This is accomplished by the F-statistic given in (3).
Tweaking the Whole Model F-Test
The full model in (1) is still called the full model. The intercept-only model is now called the reduced model. It is called the reduced model because we drop the predictors from the full model. The F-statistic in (3) is re-written as follows.
-
(4)……….
RSS with subscript 0 is error sum of squares of the reduced model, which is the TSS in (3). RSS with subscript 1 is the error sum of squares of the full model, which is the RSS in (3). In (4), is the number predictors dropped from the full model. The null hypothesis tested by the F-statistic in (4) is still the same as the one in (2), but it can be restated that the full model does not provide a significantly better fit than the reduced model. Under the null hypothesis, the F-statistic (4) will follow an F-distribution with degrees of freedom
and
.
The Generalized F-Test
With (4) as preparation, we are now ready to expand the F-test to handle “full versus reduced” and not just “full versus intercept-only.” First, we describe precisely the two models in question.
-
(5)…
(5)…(Full Model)
(5)…
(5)…(Reduced Model)
We let so that the full model in (5) has the same number of predictors as the full model in (1). Here, we drop
many predictors from the full model to form the reduced model. For clarity, we arrange the dropped predictors in the end of the model equation of the full model. The predictors in the reduced model are
. The predictors in the full model but not in the reduced model are
. The following are the hypotheses to be tested.
-
(6)……….
(6)………. for at least one
The coefficients identified in the null hypothesis are the ones for the predictors dropped from the full model. Thus,
corresponds to the reduced model. Thus, not rejecting the null hypothesis means that we choose the reduced model. On the other hand, rejecting the null hypothesis means that the predictors
as a whole should not be dropped from the full model, i.e., one of these predictors plays an important role in explaining the variability in the response. We use the following F-statistic to test the hypotheses in (6).
-
(7)……….
where is the sample size,
is the number of predictors in the full model with
and
is the number of predictors dropped from the full model. Furthermore, RSS with subscript 0 is the RSS (the error sum of squares) of the reduced model and RSS with subscript 1 is the RSS (error sum of squares) of the full model. The following clarifies these error sums of squares.
The difference is the extra sum of squares. As noted above, having more predictors in the model lowers its RSS. As a result the difference
. Here’s what we can say about the extra error sum of squares
A large extra sum of squares indicates that the predictors not found in the reduced model (or the predictors dropped from the full model) improve the prediction of the response over that obtained by regressing the response on the predictors in the reduced model. It follows from the preceding paragraph that the extra sum of squares
can be used as a basis for testing the importance of the predictors
in the presence of
. In other words,
can help us determine whether the additional predictors
can add significant predictive power over the predictors
that are already in the reduced model.
Thus, the test statistic for testing the null hypothesis in (6) should be a function of the extra sum of squares , with the direction that the larger the extra sum of squares, the larger the test statistic. The F-statistic in (7) is one such test statistics. It can be used to judge whether the extra sum of squares
is large enough to warrant the rejection of the null hypothesis.
Under the null hypothesis in (6), it can be shown that the F-statistic in (7) follows an F-distribution with degrees of freedom
and
. How large does the extra sum of squares
for us to be willing to reject the null hypothesis? The answer lies in evaluating the magnitude of the F-statistic, as described in the algorithm below.
The Algorithm
The denominator of the F-statistic in (7) is the mean squared error (MSE) of the full model. To calculate the F-statistic, we need to fit two regression models.
- Run the full model as indicated by the model equation in (5). Obtain the error sum of squares
and the MSE, i.e.,
.
- Run the reduced model as indicated by the model equation in (5). Obtain the error sum of squares
.
- Compute the F-statistic according to (7).
- Calculate the p-value under the assumption that the null hypothesis
in (6) is true, calculate the p-value under the assumption that the F-statistic in (7) follows an F-distribution with degrees of freedom
and
. Let
denote the cumulative distribution for this F-distribution. Recall that the p-value is the probability of obtaining an observed value of the F-statistic more extreme than the on observed from the data. It follows that the p-value is
where
is the observed value of the F-statistic.
- If the p-value is sufficiently small (e.g., 0.05 or 0.01), the F-statistic is large enough to warrant the rejection of the null hypothesis. We can then conclude that we should not drop the predictors
from the full model. The small p-value provides evidence that these predictors play an important role in explaining the variability in the response.
- If the p-value is not below a pre-specified threshold (e.g., 0.05 or 0.01), the value of the F-statistic is not large enough to warrant the rejection of the null hypothesis. Adding the predictors
to the reduced model does not add sufficient predictive power to the model. In other words, we conclude that the additional predictor variables not in the reduced model do not improve the prediction of the response over that obtained by regressing the response on the predictors in the reduced model.
ANOVA Tables
As indicated in the preceding section, the generalized F-test requires fitting two separate linear regression models, one for the full model and one for the reduced model. We can use ANOVA tables to organize the information in such an F-test, one table for the full model and one for the reduced model.
ANOVA Table – Full Model
| Source of Variation…. | Sum of Squares….. | df…………….. | Mean Square. | |
|---|---|---|---|---|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
ANOVA Table – Reduced Model
| Source of Variation…. | Sum of Squares….. | df…………….. | Mean Square. | |
|---|---|---|---|---|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
In each of the tables, only the items enclosed with red borders are needed for the F-test. In the table for the full model, we need the error sum of squares (the RSS with subscript 1) and its corresponding mean square. In the table for the reduced model, we need the error sum of squares, the RSS with subscript 0. Then we can compute the F-statistic according to (7).
Example 1
We work some examples using the Boston dataset from the MASS package in R. The Boston dataset contains housing data in Boston MA and the surrounding area that were collected by the US Census. It is a relatively small dataset with only 506 records. Among the 14 variables in the dataset, the response variable is medv, the median value of owner-occupied homes in $1000’s. See here for more on the dataset and the R code. The following is the linear regression that is the full model (with 13 predictors).
All the predictors are statistically significant with the except of indus and age where indus is the proportion of non-retail business acres per town and age is the proportion of owner-occupied units built prior to 1940. Let’s drop these 2 predictors from the full model. We then fit a linear model with 11 predictors (the 13 predictors from the full model excluding indus and age). The ANOVA tables are shown below.
Example 1 – ANOVA Table – Full Model – Boston Dataset
| Source of Variation…. | Sum of Squares….. | df…………….. | Mean Square. | |
|---|---|---|---|---|
|
Regression |
31,637.51 |
13 |
2,433.655 |
|
|
Error |
11,078.78 |
492 |
22.51785 |
|
|
Total |
42,716.3 |
505 |
Example 1 – ANOVA Table – Reduced Model – Boston Dataset
| Source of Variation…. | Sum of Squares….. | df…………….. | Mean Square. | |
|---|---|---|---|---|
|
Regression |
31,634.93 |
11 |
2875.903 |
|
|
Error |
11,081.36 |
494 |
22.43191 |
|
|
Total |
42,716.3 |
505 |
Using the numbers enclosed by the red borders, the following is the calculation of the F-statistic.
-
(7)……….
The value of the F-statistic is very small (close to 0)! This means that the extra sum of squares is too small. Dropping the two predictors indus and age does not deteriorate the model fit in any noticeable way. Observe that the two errors sums of squares are very close together (11,078.78 and 11,081.36). As a result, we will not reject the null hypothesis that the two dropped variables have predictive power. This implies that we should use the reduced model instead of the full model, i.e., it is OK to drop the two variables in question. In other words, we can conclude that the additional predictor variables not found in the reduced model do not improve the prediction of the response over that obtained by regressing the response on the predictors in the reduced model.
In this case, the value of the F-statistic is so small that there is no need to obtain the p-value. Just for the record, the p-value is 0.9819827, which obviously does not support the rejection of the null hypothesis.
Example 2
Example 1 shows that the predictors indus and age can be safely dropped without losing much predictive power. Now we use the reduced model in Example 1 as the full model. For the reduced model, we drop the predictor chas, which is the Charles River dummy variable (1 if tract bounds river; 0 otherwise). We use the F-test to determine whether chas is an important predictor of median house price. The two regression models are displayed in the following two ANOVA tables.
Example 2 – ANOVA Table – Full Model – Boston Dataset
| Source of Variation…. | Sum of Squares….. | df…………….. | Mean Square. | |
|---|---|---|---|---|
|
Regression |
31,634.93 |
11 |
2875.903 |
|
|
Error |
11,081.36 |
494 |
22.43191 |
|
|
Total |
42,716.3 |
505 |
Example 2 – ANOVA Table – Reduced Model – Boston Dataset
| Source of Variation…. | Sum of Squares….. | df…………….. | Mean Square. | |
|---|---|---|---|---|
|
Regression |
31,407.72 |
10 |
3140.772 |
|
|
Error |
11308.58 |
495 |
22.84561 |
|
|
Total |
42,716.3 |
505 |
Using the numbers enclosed by the red borders, the following is the calculation of the F-statistic.
-
(7)……….
The F-statistic appears to be substantially larger than 1. The p-value is 0.001551469, which is indeed small. The small p-value provides evidence to support the rejection of the null hypothesis. This implies that the predictor chas does indeed posses substantial predictive power even in the presence of all the predictors and should not be dropped from the full model. Thus, we can conclude that the additional predictor variable chas does improve the prediction of the response over that obtained by regressing the response on the predictors in the reduced model.
t-Test versus F-Test
The above discussion casts the F-test as a comparison of two nested models, i.e., the full model with all the predictors and a smaller model with a subset of the predictors. The t-test can also be viewed in such a framework. The t-test is to decide on these on these hypotheses (see here).
-
(8)……….
(8)………. ….where
is fixed with
The null hypothesis in (8) is evaluated using the t-statistic as described below.
-
(9)……….
Under the null hypothesis in (8), the t-statistic follows a t-distribution with degrees of freedom . The magnitude of the t-statistic is used to evaluate the significance of the coefficient
(see here). Even though this test is to test the significance of one parameter, the test is performed in the presence of all the other parameters (or in the presence of all the other predictors). In other words, this test is to determine the significance of one predictor in the presence of all other predictors. This is a test of the full model versus the reduced model with one and only one predictor missing. Thus, the t-test as described in (8) and (9) is to choose one of the following models.
-
(10)……….
(10)……….…………..(Reduced Model)
The t-test as described in (8) and (9) is equivalent to the F-test with the full model and reduced model described in (10). When calculating the F-statistic according to (7), since we are only dropping one predictor from the full model. In fact, the F-statistic in this case is the square of the t-statistic in (9). In Example 2, the reduced model is the result of dropping the variable chas. The value of the F-statistic is 10.129. As a result, the t-statistic is 3.183. The corresponding p-value is 0.001551, which is identical to the p-value if F-test is used.
To interpret the t-test, as in the F-test, we must do so in the presence of the other predictor. If the null hypothesis is rejected, we conclude that the inclusion of the predictor
in the model will significantly improves the prediction of the response over that obtained by regressing on all the predictors excluding
. If
is not rejected, we conclude that including the predictor
does not significantly improve the prediction of the response over that obtained by regressing the response on all the predictors excluding
.
The F-Statistic in Another Form
The F-statistic in (7) is defined using the error sum of squares in two models. It can be restated using the statistics of the full model and reduced model.
-
(11)……….
where is the
statistic for the full model and
is the
statistic for the reduced model. (11) is derived by dividing the numerator and denominator of (7) by the TSS. Adding more predictors to a model always increases the
statistic. Thus,
. The F-statistic (11) will help us determine whether the increase in
is significant when adding predictors into the reduced model.
……….
This article was originally published here in a companion site.
Dan Ma Generalized F-Test
Daniel Ma Generalized F-Test
Dan Ma F-Test
Daniel Ma F-Test
Dan Ma full model versus reduced model
Daniel Ma full model versus reduced model
Dan Ma linear regression
Daniel Ma linear regression
Dan Ma simple linear regression
Daniel Ma simple linear regression
Dan Ma multiple linear regression
Daniel Ma multiple linear regression
2024 – Dan Ma
Posted: October 14, 2023











