offline taxonomic name resolution
Offline taxonomic name matching against local Darwin Core backbones, with matching done in C.
Hand it a column of messy species names. taxify cleans
them, matches them against a backbone you already have on disk, resolves
synonyms to accepted names, and returns one standardized data.frame.
Every step runs locally against a versioned snapshot, so there are no
API calls, no rate limits, and the same input gives the same output on
any machine. The matching engine is written in C through the vectra columnar engine.
library(taxify)
# first call installs the default backbone set (COL + GBIF + ITIS, ~4 GB)
taxify(c(
"Quercus robur",
"Pinus abies", # synonym, resolved to Picea abies
"Quercus robus", # typo, fuzzy-corrected to Q. robur
"Taraxacum officinale"
))The usual route for name resolution, taxize, calls out
to around twenty web services (NCBI, ITIS, GBIF, EOL, IUCN, WoRMS,
Tropicos, …). That covers everything, but it ties each run to network
latency, service uptime, and rate limits, and the answer can change
between runs as upstream services update. taxify ships the
backbones as pre-built local snapshots and matches against them in C, so
a list of thousands resolves in seconds and a result is reproducible
from the recorded backbone version.
The closest local analogue is taxadb, which also stores backbone snapshots on disk; the migration vignette walks through the differences in matching strategy, output schema, and enrichment.
taxify ships fifteen backbones as compressed
.vtr files, pre-built by the companion taxifydb package,
downloaded once, and matched locally. Pass several and they form a
fallback chain: a name unmatched by the first backbone cascades to the
next.
# WFO first (plants), then GBIF for whatever WFO doesn't cover
taxify(
c("Quercus robur", "Panthera leo", "Amanita muscaria"),
backbone = c("wfo", "gbif")
)| Backbone | Scope | Names | Download |
|---|---|---|---|
| WFO | Vascular plants | 1.6M | 797 MB |
| COL | All kingdoms | 5.3M | 2.0 GB |
| GBIF | All kingdoms | 6.4M | 1.9 GB |
| ITIS | US focus, freshwater/marine | 992k | 205 MB |
| NCBI | All life | 2.8M | 514 MB |
| OTT | All life (synthetic) | 3.7M | 727 MB |
| WoRMS | Marine/aquatic | 1.6M | 347 MB |
| Euro+Med | European/Mediterranean plants | 147k | 35 MB |
| Species Fungorum | Fungi | 315k | 71 MB |
| AlgaeBase | Algae | 172k | 36 MB |
| FishBase | Fishes | 103k | 19 MB |
| SeaLifeBase | Non-fish marine/aquatic | 134k | 29 MB |
| Reptile Database | Reptiles | 50k | 10 MB |
| LCVP | Vascular plants | 1.3M | 252 MB |
| WCVP | Vascular plants | 1.4M | 334 MB |
list_backbones() returns this table live, with the
installed and version status of each. taxify_databases()
adds the enrichment layers alongside it for the full picture.
Input names are normalized first, so the fuzzy pass only runs on names that genuinely differ from the backbone rather than on names that just carry extra authorship or qualifiers:
"Quercus robur L." -> "Quercus robur" # authorship stripped
"Pinus cf. sylvestris" -> "Pinus sylvestris" # qualifier removed
"Nothofagus x alpina" -> "Nothofagus × alpina" # hybrid sign normalized (x -> ×)
"Betula pendula (Roth) Doll" -> "Betula pendula" # parenthesized author strippedFuzzy matching is configurable (Damerau-Levenshtein, Levenshtein, or Jaro-Winkler, with a distance threshold), and runs genus-blocked so a typo only competes against names in the same genus.
Both packages can read the same WFO snapshot, so the comparison below
is between the two matching implementations rather than between two
versions of WFO. taxify queries the columnar .vtr taxifydb
builds from the Zenodo Darwin Core archive; WorldFlora reads the
classification table inside that same archive into RAM.
The corpus is 1,000 accepted binomials drawn from the backbone with a fixed seed. The fuzzy corpus is those names with one substituted character in each epithet, so every name has to be resolved by distance – a deliberate worst case, not a mixed list where most names still match exactly.
| taxify | WorldFlora | |
|---|---|---|
| Backbone load | 4.9 s | 20.1 s (CSV into RAM) |
| Exact match, 1,000 names | 2.2 s | 17.1 s |
| Fuzzy match, 1,000 names | 18.8 s | 4,192 s (70 min) |
| Fuzzy match, 5,000 names | 26.6 s | not measured |
| Peak R heap, fuzzy 1,000 | 678 MB | 4.0 GB |
scripts/benchmark-worldflora.R produces these numbers
and scripts/benchmark-worldflora-results.json records the
run, including package versions and the backbone snapshot. Both packages
were measured back to back on one machine (Windows 11, R 4.6.0, taxify
0.3.21, WorldFlora 1.14.5) that was carrying other work at the time, so
treat the ratios as the reliable figures and the absolute times as an
upper bound.
Where the gap comes from: taxify runs the fuzzy pass genus-blocked, so a typo competes only against names in the same genus, and the comparison itself runs in C. Scaling five times the names costs 1.4x the time, because the fixed cost of opening the backbone dominates.
taxize is not in the table. It resolves names through the Global Names online resolver rather than a local backbone, so its time is network-bound and varies with service load; there is no way to measure it against a fixed snapshot the way these two can be measured against each other.
taxify() returns one row per input name with a fixed
schema: the matched and accepted names with their IDs and authorship,
rank, family, genus, epithet, synonym / hybrid / ambiguity flags, any
taxonomic qualifier, the match type (exact,
exact_ci, fuzzy, abbrev,
hybrid_formula, out_of_scope, or
none), the fuzzy distance, a coarse kingdom / taxon-group
label, the backbone, and the backbone version used.
summary() prints a compact digest of how the batch
resolved.
result <- taxify(c("Quercus robur", "Pinus abies", "Quercus robus", "Taraxacum officinale"))
summary(result)
#> -- taxify results ----------------------------------------------------
#> backbone: WFO | 4 names submitted
#>
#> matched 4 (exact: 2, case-insensitive: 0, fuzzy: 2, abbrev: 0)
#> --------------------------------------------------------------
#> taxon groups: vascular plant: 4taxify() resolves a name to its accepted name. The same
local backbone file answers the related lookups too — synonyms,
children, ancestors, descendants — with nothing else to download:
synonyms("Picea abies") # every synonym of an accepted name
children("Quercus") # accepted species in a genus
downstream("Fagaceae", downto = "genus") # all genera under a family
upstream("Quercus robur", to = "family") # the family a species sits inDecompose a name without matching it, or go from a backbone ID back to a name:
parse_name("Quercus robur (L.) H.Karst.") # genus / epithet / author / rank, no lookup
id2name("2878688", backbone = "gbif") # GBIF usage key -> name + classificationCross the vernacular boundary in either direction, and shape a lineage into a tree:
comm2sci("pedunculate oak") # common name -> scientific
sci2comm("Quercus robur") # scientific -> common names
class2tree(species) # taxonomy as a Newick / ape phylo tree
lowest_common(species) # deepest shared rank (the MRCA)For migrating an old checklist and recording exactly what you matched against:
reconcile(old_species_list) # how each name maps onto the current backbone
lock <- taxify_lock(result) # freeze the resolved backbone + enrichment versions
cite(result) # formatted citations for every source usedNearly ninety enrichment layers join published trait and status data to your results through the backbone-resolved accepted name, so synonyms in either dataset land on the same key:
# plants
taxify(plant_names) |>
add_iucn() |> # IUCN Red List
add_griis("AT") |> # GRIIS
add_zanne() |> # Zanne et al. woodiness
add_eive() # EIVE indicator values
# fish
taxify(fish_names, backbone = "col") |>
add_fishbase() |> # FishBase morphology & ecology
add_fishmorph() # FISHMORPH functional traits
# one trait, gathered across every source that has it
taxify(plant_names) |>
add_trait("seed_mass") # Diaz + GIFT, harmonized to mgSources span all kingdoms: IUCN, GRIIS, GBIF common names, WCVP,
EIVE, Diaz et al., LEDA, GIFT, FungalTraits, FUNGuild, AlgaeTraits,
EltonTraits, AVONET, PanTHERIA, AmphiBIO, FISHMORPH, FishBase, AnAge,
GloNAF, LepTraits, AnimalTraits, and regional plant-trait sets for
France (Baseflor), Britain (Ecoflora), and Germany (FloraWeb), and more.
The enrichments
vignette lists the full set with references and licenses, and
list_enrichments() returns it in R;
list_traits() browses the cross-source trait vocabulary
that add_trait() draws on.
To join your own table, add_data() auto-detects the
species column, matches it through the same backbone(s) used in the
original call, and left-joins. It accepts data.frames, CSV, CSV.GZ,
XLSX, SQLite, and .vtr.
result |> add_data("TRY_traits.csv")
result |> add_data("TRY_traits.csv", cols = c("LeafArea", "SLA", "PlantHeight"))inspect() returns only the names that look wrong, each
labelled with what stands out and the name to use instead. It catches
typos, retired synonyms, made-up genera, near-duplicate spellings, and
the lone animal in a list of plants, and ranks them by whether they need
a decision, a second look, or just optional cleanup.
inspect(field_names) # offline register and list checks
inspect(field_names, backbones = TRUE) # also typos, synonyms, ambiguityFor regional field lists, region steers a fuzzy
correction toward species that actually occur where you work, so a
misspelling resolves to the plant that grows there even when a
one-letter neighbour from another continent sits just as close in
spelling. Pass a region name, a TDWG code, or coordinates.
taxify(field_names, region = "Belgium")
taxify(field_names, coords = c(4.35, 50.85))install.packages("taxify")Or the development version from GitHub:
install.packages("pak")
pak::pak("gcol33/taxify") # vectra is installed automatically“Software is like sex: it’s better when it’s free.” — Linus Torvalds
I’m a PhD student who builds R packages in my free time because I believe good tools should be free and open. I started these projects for my own work and figured others might find them useful too.
If this package saved you some time, buying me a coffee is a nice way to say thanks. It helps with my coffee addiction.
MIT (see the LICENSE file)
@software{taxify,
author = {Colling, Gilles},
title = {taxify: Offline Taxonomic Name Matching Against Darwin Core Backbones},
year = {2026},
url = {https://github.com/gcol33/taxify}
}Cite the backbones and enrichment layers you actually used with
cite(result), which pulls each source’s own reference from
the manifest.