Parses a model formula, builds the model frame and classifies the resulting
design into one of seven dependency structures. The pieces of the design are
returned under a fixed set of names, so that every function offering a
formula interface can share one entry point instead of re-implementing the
parsing, the subset handling and the distinction between a grouping
factor, a numeric predictor and a blocking variable.
resolveFormulaFromCall() is the entry point for a function offering a
formula interface: called from its body, it forwards the caller's
formula, data and subset to resolveFormula() exactly as they were
written, so that subset keeps its base R semantics (see Details).
Usage
resolveFormula(
formula,
data,
subset,
na.action = na.pass,
allowed = c("one-sample", "two-sample-independent", "two-sample-dependent",
"n-sample-independent", "n-sample-dependent", "numeric-numeric", "regression")
)
resolveFormulaFromCall(allowed, na.action = na.pass)Arguments
- formula
a two-sided model formula. Supported forms are:
y ~ 1one-sample design.
Pair(x, y) ~ 1two-sample dependent (paired).
Pair()constructs a two-column matrix of paired observations.y ~ gtwo-sample or n-sample independent group comparison.
y ~ a:bindependent group comparison of the cells of several grouping variables, combined into one grouping factor.
y ~ a + bis not a grouped design: it isregressionif allowed, else an error.y ~ x,xnumericnumeric-numeric (correlation, simple regression).
y ~ x1 + x2 + ...general regression.
y ~ trt | blockn-sample dependent (blocked design).
- data
an optional data frame containing the variables in
formula. A matrix is coerced to a data frame.- subset
an optional expression indicating the observations to use, evaluated in
dataas inmodel.frame()(subset = len > 10), or an index vector. See Details.- na.action
a function specifying how missing values are handled, defaults to
na.pass().- allowed
a character vector restricting which design types are accepted, any combination of
"one-sample","two-sample-independent","two-sample-dependent","n-sample-independent","n-sample-dependent","numeric-numeric"and"regression". The values are matched exactly, an unknown one is an error rather than being ignored. A further error is raised if the detected type is not among the allowed ones. Defaults to all types.
Value
a named list containing at least:
- type
character, one of the design types listed above.
- mf
the
model.frame()the design was read from.- rows
integer, the positions of the retained observations in the original data, or
NULLif they cannot be determined.- response
the left-hand side of the formula.
- dataName
character, the deparsed formula, for use as the
data.nameof anhtestobject.
plus the design-specific components described under Details.
Details
Design types
one-sampley ~ 1, as in the one-sample t-test or the one-sample Wilcoxon test.two-sample-independenty ~ gwith two groups, as in the two-sample t-test or the Wilcoxon rank-sum test.two-sample-dependentPair(x, y) ~ 1, as in the paired t-test or the Wilcoxon signed-rank test.n-sample-independenty ~ gwith more than two groups, as in the analysis of variance or the Kruskal-Wallis test.n-sample-dependenty ~ trt | block, as in a repeated-measures analysis of variance or the Friedman test.numeric-numericy ~ xwith a numeric right-hand side, as in correlation or simple regression.regressiona general regression formula with one or more predictors.
Type detection
The type follows from the shape of the formula and from the class of the
right-hand side variable, not from allowed. allowed only decides
whether the detected type is accepted, with four exceptions worth knowing.
A grouping factor carrying a single level is reported as
one-sampleif that type is allowed.A two-group design is reported as
n-sample-independentif"two-sample-independent"is not among the allowed types. Together with the previous rule this lets a caller that treats every group count alike allow one type only.allowed = "regression"on its own forces theregressiontype for every formula, includingy ~ 1andy ~ g, with the exception of the blocked syntaxy ~ trt | block, which is alwaysn-sample-dependentand therefore an error then. This is the entry point for model-fitting callers, which interpret the right-hand side themselves. Duplicates inallowedare ignored, soc("regression", "regression")behaves the same.Otherwise
regressionis reported only for more than one right-hand side variable. Withallowedcontaining both"regression"and"numeric-numeric",y ~ xis thereforenumeric-numericwhiley ~ x1 + x2isregression.
Cells of several grouping variables
A right-hand side consisting of a single interaction term, y ~ a:b (or
y ~ a:b:c), is the explicit request for the cells of these variables.
They are combined into one grouping factor via interaction(), with
levels such as "OJ:0.5"; numeric components are treated as categorical.
The design is then classified like y ~ g, by the number of non-empty
cells.
Unlike boxplot(), y ~ a + b is not read as cells: additive terms are
not an interaction, as in lm(). Such a formula, like y ~ a * b, is
regression if that type is allowed, and an error otherwise. If
"regression" is allowed, it also takes precedence over the cell reading
of y ~ a:b, which a model-fitting caller interprets as an interaction
term.
The distinction rests on the terms of the formula: the model frame
holds the variables a and b in both cases and cannot tell the two
apart.
offset() is only accepted for the regression design and an error
otherwise: an offset column is part of the model frame without being a
term, and would pass for a grouping variable or a numeric predictor.
Field naming contract (binding across all types)
responseis present for every type and always holds the left-hand side of the formula. It is the one field a caller can rely on without branching ontype.xis an alias ofresponsefor the types that are conventionally described in terms of a sample rather than a model (one-sample,two-sample-independent,n-sample-independent,numeric-numeric). The exception istwo-sample-dependent: thereresponseis the wholePair()matrix, whilexandyare its first and second column.groupis reserved for a categorical, factor-coercible variable of lengthn(the full sample) that splits the response into groups. It is never pre-split and never used for a continuous variable.xandgrouphave an identical shape fortwo-sample-independentand forn-sample-independent, so that a caller can usesplit(r$x, r$group)uniformly, without branching on the number of groups. Fory ~ a:b,groupholds the combined cell factor.predictoris used for a continuous, numeric right-hand side variable (numeric-numeric), nevergroup.treatmentis used for the explanatory variable of a blocked design (n-sample-dependent), as distinct fromblock, the stratification factor. Neither is ever calledgroup.y, where present, is a convenience field only, holding the second group of a two-sample design or the second paired vector. It is never needed for correct use:xandgroup(orxandpredictor, ortreatmentandblock) are always sufficient and are the canonical access path.rowsis present for every type and holds the positions of the retained observations in the original data, aftersubsetand afterna.action. It is the handle for synchronising an external vector (an ordering variable, weights) with the model frame:z <- z[r$rows]. It isNULLin the rare case where the row names of the model frame cannot be matched back, e.g. when a numericsubsetselects a row twice.termsis returned for theregressiontype. Build the design matrix from it,model.matrix(r$terms, r$mf), never from the original formula: the columns of a model frame are named after the deparsed expressions ("log(x)"), so re-evaluating the formula against the model frame fails for every transformed term.
Missing values
Missing values are left to na.action and are not touched otherwise, so
with the default na.pass() they reach the caller untouched. The one
exception is the grouping factor of an independent design, where empty and
missing levels are dropped before the groups are counted. A grouping
variable that is missing throughout leaves no level at all and is an error.
A cell of y ~ a:b is missing as soon as one of its components is.
Rows removed by na.action are recorded in attr(r$mf, "na.action"), but
those indices are relative to the already subsetted frame. To align an
external vector with the model frame use rows, which accounts for
subset and na.action at once.
subset handling
subset is evaluated as in base R: it is taken unevaluated, like
lm() does, by handing this function's own call on to model.frame(),
which evaluates the expression in data, with environment(formula) as
the enclosure. resolveFormula(y ~ g, df, subset = g == "A") therefore
works exactly like boxplot(y ~ g, df, subset = g == "A"), and a quoted
expression fails in both.
A wrapper function forwards its arguments the same way base R does (see
boxplot.formula()): it rebuilds its own call and evaluates it in its
caller's frame, so that subset reaches resolveFormula() unevaluated.
resolveFormulaFromCall() does exactly this and is the entry point to use
in a formula method:
myFun <- function(formula, data, subset, na.action = na.omit, ...) {
r <- resolveFormulaFromCall(
allowed = c("two-sample-independent", "n-sample-independent"),
na.action = na.action)
...
}Passing a captured expression on by value
(resolveFormula(formula, data, subset = substitute(subset))) does not
work, just as it does not for model.frame() itself.
Return components by type
Every return value starts with type, mf, rows and response, in that
order, followed by the design-specific components below, and ends with
dataName:
one-samplextwo-sample-independentx,group,y(convenience: the second group)two-sample-dependentx,yn-sample-independentx,groupn-sample-dependenttreatment,blocknumeric-numericx,predictorregressionterms
resolveFormulaFromCall()
Rebuilds the call of the function it is called from - via match.call()
against that function's definition, which works for an S3 method as well
keeps its
formula,dataandsubsetarguments, and evaluatesresolveFormula()with them in the frame the calling function was called from. This is the forwarding pattern ofboxplot.formula()andlm(), written once:
plotBox.formula <- function(formula, data, subset, na.action = na.omit, ...) {
r <- bedrock::resolveFormulaFromCall(
allowed = c("two-sample-independent", "n-sample-independent"),
na.action = na.action)
...
}allowed is passed on as given (omitted: all types). na.action should
be the calling function's own na.action value; it is always passed on,
so that the caller's default (typically na.omit()) applies rather than
the na.pass() default of resolveFormula().
Requirements for the calling function: its arguments must be named
formula, data and subset, and resolveFormulaFromCall() must be
called directly in its body, not from a nested helper function, since the
call is looked up one frame up.
See also
model.frame(), Pair(), resolveGroups()
Other data.resolve:
resolveContingency(),
resolveGroups()
Examples
set.seed(1)
df <- data.frame(
y = rnorm(30, 50, 10),
g2 = rep(c("A", "B"), 15),
g3 = rep(c("A", "B", "C"), 10),
trt = rep(c("T1", "T2", "T3"), 10),
blk = rep(1:10, 3)
)
# one-sample
resolveFormula(y ~ 1, data = df)$type
#> [1] "one-sample"
## [1] "one-sample"
# two-sample independent: x and group have full length, the same shape
# as for more than two groups
r2 <- resolveFormula(y ~ g2, data = df,
allowed = c("two-sample-independent",
"n-sample-independent"))
r2$type
#> [1] "two-sample-independent"
## [1] "two-sample-independent"
length(r2$x) == length(r2$group)
#> [1] TRUE
## [1] TRUE
# n-sample independent
resolveFormula(y ~ g3, data = df,
allowed = "n-sample-independent")$type
#> [1] "n-sample-independent"
## [1] "n-sample-independent"
# cells of two grouping variables: a:b, not a + b
r3 <- resolveFormula(y ~ g2:g3, data = df,
allowed = "n-sample-independent")
levels(r3$group)
#> [1] "A:A" "B:A" "A:B" "B:B" "A:C" "B:C"
## [1] "A:A" "B:A" "A:B" "B:B" "A:C" "B:C"
try(resolveFormula(y ~ g2 + g3, data = df,
allowed = "n-sample-independent"))
#> Error : 'formula' should be of the form response ~ group; use response ~ a:b for the cells of several grouping variables
# two-sample dependent (paired)
df2 <- data.frame(pre = rnorm(15, 50, 10), post = rnorm(15, 55, 10))
resolveFormula(Pair(pre, post) ~ 1, data = df2,
allowed = c("one-sample",
"two-sample-dependent"))$type
#> [1] "two-sample-dependent"
## [1] "two-sample-dependent"
# n-sample dependent (blocked): treatment, not group
r4 <- resolveFormula(y ~ trt | blk, data = df,
allowed = "n-sample-dependent")
names(r4)
#> [1] "type" "mf" "rows" "response" "treatment" "block"
#> [7] "dataName"
## [1] "type" "mf" "rows" "response" "treatment" "block" "dataName"
# numeric-numeric: predictor, not group
df3 <- data.frame(y = rnorm(20), x = rnorm(20))
r5 <- resolveFormula(y ~ x, data = df3, allowed = "numeric-numeric")
is.numeric(r5$predictor)
#> [1] TRUE
## [1] TRUE
# regression: build the design matrix from 'terms', not from the formula
r6 <- resolveFormula(y ~ log(abs(x)) + I(x^2), data = df3,
allowed = "regression")
colnames(model.matrix(r6$terms, r6$mf))
#> [1] "(Intercept)" "log(abs(x))" "I(x^2)"
# subset, as in base R
resolveFormula(y ~ g3, data = df, subset = g3 != "C",
allowed = "two-sample-independent")$type
#> [1] "two-sample-independent"
## [1] "two-sample-independent"
