| Type: | Package |
| Title: | Accessing the 'Wordbank' Database |
| Description: | Connecting to 'Wordbank' https://wordbank.stanford.edu, an open repository for developmental vocabulary data from the MacArthur-Bates Communicative Development Inventories (Frank et al. 2017 <doi:10.1017/S0305000916000209>). Data are read from a versioned dataset hosted on 'Redivis', so analyses can be pinned to a specific release. |
| Version: | 2.0.0 |
| Depends: | R (≥ 4.1.0) |
| License: | GPL-3 |
| URL: | https://langcog.github.io/wordbankr/, https://github.com/langcog/wordbankr/ |
| BugReports: | https://github.com/langcog/wordbankr/issues/ |
| Imports: | assertthat (≥ 0.2.1), dplyr (≥ 1.1.3), glue (≥ 1.6.2), jsonlite (≥ 1.8.7), lifecycle (≥ 1.0.3), purrr (≥ 1.0.2), quantregGrowth (≥ 1.7-0), rlang (≥ 1.1.1), robustbase (≥ 0.99-0), stringr (≥ 1.5.0), tidyr (≥ 1.3.0) |
| Suggests: | redivis (≥ 0.12.12), ggplot2, knitr, rmarkdown, testthat (≥ 3.0.0) |
| Additional_repositories: | https://langcog.r-universe.dev |
| VignetteBuilder: | knitr |
| Config/testthat/edition: | 3 |
| Encoding: | UTF-8 |
| Config/roxygen2/version: | 8.1.0 |
| RoxygenNote: | 7.3.3 |
| NeedsCompilation: | no |
| Packaged: | 2026-09-30 16:55:10 UTC; mcfrank |
| Author: | Mika Braginsky [aut, cre, cph], Daniel Yurovsky [ctb], Michael Frank [ctb], Danielle Kellier [ctb], Alvin Tan [ctb] |
| Maintainer: | Mika Braginsky <mika.br@gmail.com> |
| Repository: | CRAN |
| Date/Publication: | 2026-10-10 11:00:02 UTC |
wordbankr: Accessing the 'Wordbank' Database
Description
Connecting to 'Wordbank' https://wordbank.stanford.edu, an open repository for developmental vocabulary data from the MacArthur-Bates Communicative Development Inventories (Frank et al. 2017 doi:10.1017/S0305000916000209). Data are read from a versioned dataset hosted on 'Redivis', so analyses can be pinned to a specific release.
Author(s)
Maintainer: Mika Braginsky mika.br@gmail.com [copyright holder]
Other contributors:
Daniel Yurovsky yurovsky@stanford.edu [contributor]
Michael Frank mcfrank@stanford.edu [contributor]
Danielle Kellier dkellier@stanford.edu [contributor]
Alvin Tan tanawm@stanford.edu [contributor]
See Also
Useful links:
Report bugs at https://github.com/langcog/wordbankr/issues/
Deprecated database arguments
Description
As of wordbankr 2.0, data are retrieved from the versioned Wordbank dataset on Redivis rather than a MySQL database; 'db_args' and 'connect_to_wordbank()' are deprecated and ignored.
Usage
check_db_args(db_args)
connect_to_wordbank(db_args = NULL)
get_wordbank_args()
Arguments
db_args |
Deprecated, ignored. |
Value
check_db_args(): no return value, called for its side
effect (a warning if db_args is supplied).
connect_to_wordbank(): the Wordbank dataset reference, as
returned by wb_dataset().
get_wordbank_args(): a list with elements organization,
dataset, and version identifying the Redivis dataset.
Fit age of acquisition estimates for Wordbank data
Description
For each item in the input data, estimate its age of acquisition as the earliest age (in months) at which the proportion of children who understand/produce the item is greater than some threshold. The proportions used can be empirical or first smoothed by a model.
Usage
fit_aoa(
instrument_data,
measure = "produces",
method = "glm",
proportion = 0.5,
age_min = min(instrument_data$age, na.rm = TRUE),
age_max = max(instrument_data$age, na.rm = TRUE)
)
Arguments
instrument_data |
A data frame returned by |
measure |
One of "produces" or "understands" (defaults to "produces"). |
method |
A string indicating which smoothing method to use:
|
proportion |
A number between 0 and 1 indicating threshold proportion of children. |
age_min |
The minimum age to allow for an age of acquisition. Defaults
to the minimum age in |
age_max |
The maximum age to allow for an age of acquisition. Defaults
to the maximum age in |
Value
A data frame where every row is an item, the item-level columns from
the input data are preserved, and the aoa column contains the age of
acquisition estimates.
Examples
eng_ws_data <- get_instrument_data(language = "English (American)",
form = "WS",
items = c("item_1", "item_42"),
administration_info = TRUE)
if (!is.null(eng_ws_data)) eng_ws_aoa <- fit_aoa(eng_ws_data)
Fit quantiles to vocabulary sizes using quantile regression
Description
Fit quantiles to vocabulary sizes using quantile regression
Usage
fit_vocab_quantiles(vocab_data, measure, group, quantiles = "standard")
Arguments
vocab_data |
A data frame returned by |
measure |
A column of |
group |
(Optional) A column of |
quantiles |
Either one of "standard" (default), "deciles", "quintiles", "quartiles", "median", or a numeric vector of quantile values. |
Value
A data frame with the columns "language", "form", "age", group
(if specified), "quantile", and measure, where measure is the
fit vocabulary value for that quantile at that age.
Examples
eng_wg <- get_administration_data(language = "English (American)",
form = "WG",
include_demographic_info = TRUE)
if (!is.null(eng_wg)) {
vocab_quantiles <- fit_vocab_quantiles(eng_wg, production)
vocab_quantiles_sex <- fit_vocab_quantiles(eng_wg, production, sex)
vocab_quartiles <- fit_vocab_quantiles(eng_wg, production, quantiles = "quartiles")
}
Get the Wordbank by-administration data
Description
Get the Wordbank by-administration data
Usage
get_administration_data(
language = NULL,
form = NULL,
filter_age = TRUE,
include_demographic_info = FALSE,
include_birth_info = FALSE,
include_health_conditions = FALSE,
include_language_exposure = FALSE,
include_study_internal_id = FALSE,
version = "current"
)
Arguments
language |
An optional string specifying which language's administrations to retrieve. |
form |
An optional string specifying which form's administrations to retrieve. |
filter_age |
A logical indicating whether to filter the administrations to ones in the instrument's age range. |
include_demographic_info |
A logical indicating whether to include the
child's demographic information ( |
include_birth_info |
A logical indicating whether to include the
child's birth information ( |
include_health_conditions |
A logical indicating whether to include
the child's health condition information (a nested dataframe under
|
include_language_exposure |
A logical indicating whether to include
the child's language exposure information at time of administration (a
nested dataframe under |
include_study_internal_id |
A logical indicating whether to include the child's ID in the original study data. |
version |
A string specifying which version of the Wordbank dataset
to use, e.g. |
Value
A data frame where each row is a CDI administration and each column
is a variable about the administration or the corresponding child,
including which dataset_version it came from.
Examples
english_ws_admins <- get_administration_data("English (American)", "WS")
Get cached age-of-acquisition estimates
Description
Age-of-acquisition estimates for every word item on every instrument,
precomputed with fit_aoa (glm method, 50
data release. aoa is NA for items that do not reach the
threshold within the instrument's age range.
Usage
get_aoa(language = NULL, form = NULL, measure = NULL, version = "current")
Arguments
language |
An optional string specifying which language's estimates to retrieve. |
form |
An optional string specifying which form's estimates to retrieve. |
measure |
An optional string ( |
version |
A string specifying which version of the Wordbank dataset
to use, e.g. |
Value
A data frame with one row per instrument item and measure:
language, form, item_id, item_definition,
category, uni_lemma, measure, aoa,
dataset_version.
Examples
danish_aoa <- get_aoa(language = "Danish", form = "WS")
Get item-by-age summary statistics for items across languages
Description
Get item-by-age summary statistics for items across languages
Usage
get_crossling_data(uni_lemmas, version = "current")
Arguments
uni_lemmas |
A character vector of uni_lemmas. |
version |
A string specifying which version of the Wordbank dataset
to use, e.g. |
Value
A dataframe with a row for each combination of language, item, and
age, and columns for summary statistics for the group: number of children
(n_children), means (comprehension, production),
standard deviations (comprehension_sd, production_sd); and
item-level variables (item_id, definition, uni_lemma,
lexical_category, lexical_class, dataset_version).
Examples
## Not run:
# downloads item-level data from every instrument, which takes minutes
crossling_data <- get_crossling_data(uni_lemmas = "dog")
## End(Not run)
Get the uni_lemmas available in Wordbank
Description
Get the uni_lemmas available in Wordbank
Usage
get_crossling_items(version = "current")
Arguments
version |
A string specifying which version of the Wordbank dataset
to use, e.g. |
Value
A data frame with the columns uni_lemma and
dataset_version.
Examples
uni_lemmas <- get_crossling_items()
Get the Wordbank data sources
Description
Get the Wordbank data sources
Usage
get_datasets(
language = NULL,
form = NULL,
admin_data = FALSE,
version = "current"
)
Arguments
language |
An optional string specifying which language's datasets to retrieve. |
form |
An optional string specifying which form's datasets to retrieve. |
admin_data |
A logical indicating whether to include the number of administrations in the dataset. |
version |
A string specifying which version of the Wordbank dataset
to use, e.g. |
Value
A data frame where each row is a particular dataset and its
characteristics, including which dataset_version it came from.
Examples
english_ws_datasets <- get_datasets("English (American)", "WS")
Get multilingual item embeddings
Description
Semantic embeddings for every unique word item definition, computed with
Google's multilingual gemini-embedding-001 model (768 dimensions).
All languages share one embedding space, so cosine similarities are
meaningful both within and across languages.
Usage
get_embeddings(language = NULL, version = "current")
Arguments
language |
An optional string specifying which language's embeddings to retrieve. |
version |
A string specifying which version of the Wordbank dataset
to use, e.g. |
Value
A data frame with one row per unique item definition:
language, item_definition, embedding (a
list-column of numeric vectors), and dataset_version.
Examples
danish_embeddings <- get_embeddings(language = "Danish")
Get the Wordbank administration-by-item data
Description
Get the Wordbank administration-by-item data
Usage
get_instrument_data(
language,
form,
items = NULL,
administration_info = FALSE,
item_info = FALSE,
version = "current",
...
)
Arguments
language |
A string of the instrument's language. |
form |
A string of the instrument's form. |
items |
A character vector of item ids (e.g. |
administration_info |
Either a logical indicating whether to include
administration data or a data frame of administration data (as returned
by |
item_info |
Either a logical indicating whether to include item data
or a data frame of item data (as returned by |
version |
A string specifying which version of the Wordbank dataset
to use, e.g. |
... |
Additional arguments, ignored (for backward compatibility). |
Value
A data frame where each row contains the values (value,
produces, understands) of a given item (item_id) for
a given administration (data_id), with additional columns of
variables about the administration and item, as specified, and
dataset_version.
Examples
eng_ws_data <- get_instrument_data(language = "English (American)",
form = "WS",
items = c("item_1", "item_42"))
Get the Wordbank instruments
Description
Get the Wordbank instruments
Usage
get_instruments(version = "current")
Arguments
version |
A string specifying which version of the Wordbank dataset
to use, e.g. |
Value
A data frame where each row is a CDI instrument and each column is
a variable about the instrument (instrument_id, language,
form, form_type, age_min, age_max,
has_grammar, unilemma_coverage, dataset_version).
Examples
instruments <- get_instruments()
Get the Wordbank by-item data
Description
Get the Wordbank by-item data
Usage
get_item_data(language = NULL, form = NULL, version = "current")
Arguments
language |
An optional string specifying which language's items to retrieve. |
form |
An optional string specifying which form's items to retrieve. |
version |
A string specifying which version of the Wordbank dataset
to use, e.g. |
Value
A data frame where each row is a CDI item and each column is a
variable about it: item_id, item_kind,
item_definition, english_gloss, language,
form, form_type, category, lexical_category,
lexical_class, complexity_category, uni_lemma,
dataset_version.
Examples
english_ws_items <- get_item_data("English (American)", "WS")
Get item-by-age summary statistics
Description
Get item-by-age summary statistics
Usage
summarise_items(item_data, version = "current")
Arguments
item_data |
A dataframe as returned by |
version |
A string specifying which version of the Wordbank dataset
to use, e.g. |
Value
A dataframe with a row for each combination of item and age, and
columns for summary statistics for the group: number of children
(n_children), means (comprehension, production),
standard deviations (comprehension_sd, production_sd); also
retains item-level variables from lang_items (item_id,
item_definition, uni_lemma, lexical_category) and
dataset_version.
Examples
italian_items <- get_item_data(language = "Italian", form = "WG")
if (!is.null(italian_items)) {
italian_dog <- dplyr::filter(italian_items, uni_lemma == "dog")
italian_dog_summary <- summarise_items(italian_dog)
}
The Wordbank dataset on Redivis
Description
Returns a reference to the Wordbank Redivis dataset (https://stanford.redivis.com/datasets/627v-9ewzpdvz0).
Usage
wb_dataset(version = "current")
Arguments
version |
A string specifying which version of the Wordbank dataset
to use, e.g. |
Value
A redivis dataset reference.