Data Quality library for dwc:scientificName and dwc:scientificNameAuthorship and related terms
Abstracted from the FilteredPush FP-KurationServices Scientific Name Service classes.
DOI: 10.5281/zenodo.7026712
This library provides a set of methods for validating information related to scientific names, expressed as Darwin Core terms. It implements the TDWG Biodiversity Data Quality BDQ Standard tests for scientific names, and provides a set of utilities for comparing scientific name authorship strings.
The sci_name_qc library implements the following BDQ Standard tests, with these implementations passing against all test validation records.
- VALIDATION_TAXONRANK_STANDARD
- AMENDMENT_SCIENTIFICNAMEID_FROM_TAXON
- AMENDMENT_TAXONRANK_STANDARDIZED
- VALIDATION_FAMILY_FOUND
- AMENDMENT_SCIENTIFICNAME_FROM_SCIENTIFICNAMEID
- VALIDATION_KINGDOM_FOUND
- VALIDATION_SCIENTIFICNAMEID_NOTEMPTY
- VALIDATION_SCIENTIFICNAMEID_COMPLETE
- VALIDATION_KINGDOM_NOTEMPTY
- VALIDATION_SCIENTIFICNAMEAUTHORSHIP_NOTEMPTY
- VALIDATION_TAXONRANK_NOTEMPTY
- VALIDATION_POLYNOMIAL_CONSISTENT
- VALIDATION_TAXON_NOTEMPTY
- VALIDATION_ORDER_FOUND
- VALIDATION_CLASSIFICATION_CONSISTENT
- VALIDATION_PHYLUM_FOUND
- VALIDATION_SCIENTIFICNAME_NOTEMPTY
- VALIDATION_TAXON_UNAMBIGUOUS
- VALIDATION_GENUS_FOUND
- VALIDATION_NAMEPUBLISHEDINYEAR_NOTEMPTY
- VALIDATION_CLASS_FOUND
- VALIDATION_SCIENTIFICNAME_FOUND
Provides AuthorNameComparator, a small set of classes to make nomenclatural code and convention aware comparisons of scientific name authorship strings. Example use:
String authorship = "(J. C. Schmidt) Coker & Beers ex Pouzar: Fries";
AuthorNameComparator comparator = AuthorNameComparator.authorNameComparatorFactory(authorship, null);
// Botanical formulation is recognised, and comparator is an ICNafpAuthorNameComparator.
String result = comparator.compare("(J. C. Schmidt) Coker & Beers ex Pouzar: Fries","(Schmidt) Coker & Beers ex Pouzar: Fr.").getMatchType();
System.out.println(result); // Same Author, but abbreviated differently J. C. Schmidt equated with Schmidt, Fries equated with Fr..
result = comparator.compare("(J. C. Schmidt) Coker & Beers ex Pouzar: Fries","(Schmidt) Coker & Beers ex Pouzar; Fries").getMatchType();
System.out.println(result); // Authorship botanical parts tokenize differently, recognizes : to mark sanctioning author Fries isn't present.
Returns:
Same Author, but abbreviated differently.
Authorship botanical parts tokenize differently
Currently recognised comparisons:
public static final String MATCH_EXACT = "Exact Match";
public static final String MATCH_ERROR = "Error in making comparison";
public static final String MATCH_CONNECTFAILURE = "Error connecting to service";
public static final String MATCH_FUZZY_SCINAME = "Fuzzy Match on Scientific Name";
public static final String MATCH_DISSIMILAR = "Author Dissimilar";
public static final String MATCH_STRONGDISSIMILAR = "Author Strongly Dissimilar";
/**
* Zoological, "Sowerby" in one case, Sowerby I, Sowerby II, or Sowerby III in the other,
* specification of which of the Sowerby family was the author was added.
*/
public static final String MATCH_SOWERBYEXACTYEAR = "Specifying Which Sowerby, Year Exact";
/**
* Zoological, L. may be abbreviation for Linnaeus or Lamarck.
* Does not apply to Botany.
*/
public static final String MATCH_L_EXACTYEAR = "Ambiguous L., Year Exact";
/**
* Zoological, L. may be abbreviation for Linnaeus or Lamarck.
* Does not apply to Botany.
*/
public static final String MATCH_L = "Ambiguous L.";
public static final String MATCH_WEAKEXACTYEAR = "Slightly Similar Author, Year Exact";
public static final String MATCH_SIMILAREXACTYEAR = "Similar Author, Year Exact";
public static final String MATCH_SIMILARMISSINGYEAR = "Similar Author, Year Removed";
public static final String MATCH_SIMILARADDSYEAR = "Similar Author, Year Added";
public static final String MATCH_EXACTDIFFERENTYEAR = "Exact Author, Years Different";
public static final String MATCH_EXACTMISSINGYEAR = "Exact Author, Year Removed";
public static final String MATCH_EXACTADDSYEAR = "Exact Author, Year Added";
public static final String MATCH_PARENTHESIESDIFFER = "Differ only in Parenthesies";
public static final String MATCH_PARENYEARDIFFER = "Differ in Parenthesies and Year";
public static final String MATCH_AUTHSIMILAR = "Author Similar";
public static final String MATCH_ADDSAUTHOR = "Author Added";
public static final String MATCH_MULTIPLE = "Multiple Matches:";
/**
* Botanical parenthetical, ex, sanctioning author bits are composed differently, without
* regard to the similarity or difference of the individual bits. e.g. "(x) y" is
* different from "x ex y" or (x) y ex z", but not different in this comparison from "(a) b".
*/
public static final String MATCH_PARTSDIFFER = "Authorship botanical parts tokenize differently";
/**
* Authorship strings appear similar (parenthesies, botanical functional parts), but the differences
* between the strings can be explained by differences in abbreviation or addition/removal of initials.
*/
public static final String MATCH_SAMEBUTABBREVIATED = "Same Author, but abbreviated differently.";
Limited command line access is provided to process a CSV file containing taxon names, checking those names against a specified authority and asserting matches of those names against the authority.
$ java -jar sci_name_qc-1.1.3-SNAPSHOT-{commit}-executable.jar --help
usage: SciNameUtils
-f,--file <arg> Input csv file from which to lookup names. Assumes
a csv file, first three columns being dbpk,
scientificname, authorship, (TODO: family), columns
after the third are ignored.
-h,--help Print this message
-o,--output <arg> Output file into which to write results of lookup,
default output.csv
-s,--service <arg> Service to lookup names against WoRMS,
GBIF_BACKBONE, GBIF_ITIS, GBIF_FAUNA_EUROPEA,
GBIF_UKSI, GBIF_IPNI, GBIF_INDEXFUNGORUM, GBIF_COL,
GBIF_PALEOBIOLOGYDB, or ZooBank (TODO:
WoRMS+ZooBank).
-t,--test Test connectivity with an example name
Example regex (in vim) to convert csv output to sql queries to add WoRMS guids to MCZbase:
%s/"\([0-9]\+\)","[A-Za-z.() ]\+",".*","\(urn:lsid:marinespecies.org:taxname:[0-9]\+\)",".*$/update taxonomy set taxonid = '\2', taxonid_guid_type = 'WoRMS LSID' where taxonid is null and taxon_name_id = \1;/
Class: org.filteredpush.qc.sciname.DwCSciNameDQ
Implements the TDWG BDQ TG2 Scientific Name (NAME) tests.
WoRMSService and IRMNGService share a single configured HTTP client per service, which:
- identifies itself with the User-Agent
FilteredPush-sci_name_qc/{version} (+https://github.com/FilteredPush/sci_name_qc), - reuses connections (connection pooling and keep-alive) and has explicit connect, read, and write timeouts,
- limits the number of concurrent in-flight requests to each service, and the minimum interval between requests, so that many concurrent callers (e.g. multithreaded test execution) produce a throttled stream of requests rather than a burst,
- retries only plausibly transient failures (HTTP 408, 429, 500, 502, 503, 504, and connection failures), with exponential
backoff and jitter, honoring a
Retry-Afterheader, and does not retry other failures (e.g. 400, 401, 403, 404), - retries each call to the service separately (so, for example, a failed habitat lookup does not resend the search by name),
for
validate()and for the static lookup methods (lookupTaxon,lookupTaxonByID,lookupGenus,lookupTaxonAtRank,simpleNameSearch,nameComparisonSearch), which report failures as anApiExceptioncarrying the HTTP status code, - caches the results of
validate()(keyed on scientific name, authorship, and kingdom), of searches by name, and of record lookups by AphiaID/IRMNG_ID (used for habitat lookups andlookupTaxonByID), shared byvalidate()and the static lookup methods, so repeated lookups are not resent to the service, - makes a single request when several threads look up the same value at once, sharing the result between them,
- remembers a failed lookup for a period, during which the same lookup fails without being resent,
- waits a limited time for a turn to make a request, rather than blocking indefinitely behind a slow service,
- has a circuit breaker for each service: after a number of consecutive failed calls, calls fail without being sent
(with a
ServiceUnavailableException, aServiceException) for a period, then a single trial call is allowed, and calls resume if it succeeds. Responses with non-transient error statuses (e.g. 404) show the service is responding, so do not count as failures.
Failures are logged, and reported in ServiceException messages, with the HTTP status code, request URL, Retry-After and
Content-Type headers, and (truncated) response body, or with the type of connection failure. ServiceException.getHttpStatusCode()
returns the HTTP status code (or 0 for connection and parsing failures).
These settings can be changed with java system properties, or with the static setters on org.filteredpush.qc.sciname.services.ServiceClientConfig:
| System property | Default | Meaning |
|---|---|---|
sci_name_qc.userAgent |
FilteredPush-sci_name_qc/{version} (+https://github.com/FilteredPush/sci_name_qc) |
User-Agent header |
sci_name_qc.maxRetries |
3 | Retries after a transient failure (total attempts = maxRetries + 1) |
sci_name_qc.backoffBaseMillis |
500 | Base delay for exponential backoff |
sci_name_qc.backoffMaxMillis |
8000 | Maximum backoff delay |
sci_name_qc.maxRetryAfterMillis |
30000 | Longest Retry-After that will be waited for, longer requests fail without retrying |
sci_name_qc.maxConcurrentRequests |
2 | Maximum concurrent in-flight requests to each service |
sci_name_qc.minRequestIntervalMillis |
100 | Minimum interval between the start of requests to each service (0 for none) |
sci_name_qc.connectTimeoutMillis |
10000 | Connect timeout |
sci_name_qc.readTimeoutMillis |
30000 | Read timeout |
sci_name_qc.writeTimeoutMillis |
30000 | Write timeout |
sci_name_qc.cacheSize |
10000 | Maximum entries in each lookup cache, 0 disables caching |
sci_name_qc.acquireTimeoutMillis |
60000 | Longest wait for a turn to make a request, after which it fails without being sent |
sci_name_qc.failureCacheMillis |
60000 | How long a failed lookup is remembered, 0 to not remember failures |
sci_name_qc.circuitBreakerThreshold |
5 | Consecutive failed calls that trip the circuit breaker for a service, 0 to disable it |
sci_name_qc.circuitBreakerOpenMillis |
60000 | How long calls fail without being sent once the circuit breaker has tripped |
The retry, backoff, cache size, failure cache, and circuit breaker settings take effect immediately (see below for how a change to the failure cache applies to failures already remembered), the User-Agent, timeout, concurrency, acquire timeout, and request interval settings are read when the client for a service is first used, so should be set before any lookups are made, for example:
java -Dsci_name_qc.maxConcurrentRequests=4 -Dsci_name_qc.maxRetries=5 -jar ...
When a lookup fails, after any retries, the failure is remembered for sci_name_qc.failureCacheMillis (default 60000 ms, one minute).
Until then, repeating the same lookup fails immediately without sending a request to the service. This keeps a failing lookup
from being resent for every record that contains the same name.
-
What is remembered: failures after retries have been exhausted, including non-transient HTTP errors (e.g. 400, 403, 404) and responses that could not be parsed. Failures are not remembered for calls that were never sent (the circuit breaker was open, or no turn to make the request became available in time), or that were interrupted.
-
What counts as the same lookup: failures are remembered for each cached lookup separately:
validate(): the scientific name, authorship, and kingdom;- searches by name: the name and the marine only flag;
- record lookups: the AphiaID or IRMNG_ID.
Because
validate()and the static lookup methods share the searches by name, a failed search for a name also affects the other lookups of that name (e.g.validate()with a different authorship, orlookupTaxon()), until the failure expires. -
How a remembered failure is reported: as a
ServiceUnavailableException(aServiceException), with a message startingLookup failed recently, not resending for another N ms:followed by the original failure, and with the original HTTP status code fromgetHttpStatusCode(). The static lookup methods report it as anApiExceptionwith that status code. The BDQ tests report it, as they report other service failures, as EXTERNAL_PREREQUISITES_NOT_MET, with the message in the result comments.
A WoRMS 403 (a name WoRMS will not look up) is still treated as no match when it is remembered. -
The drawback: if the service recovers during the period, lookups that failed shortly before still fail until their failures expire.
Setting sci_name_qc.failureCacheMillis=0 turns this off, so every lookup that is not answered from the cache is sent to the
service again, with its retries:
java -Dsci_name_qc.failureCacheMillis=0 -jar ...
or, in code, ServiceClientConfig.setFailureCacheMillis(0L);. The change applies to failures that happen after it is made.
Failures already remembered continue to apply until they expire, or until they are cleared with WoRMSService.clearCaches() or
IRMNGService.clearCaches(), which clear both the cached results and the remembered failures for that service.
Related settings work independently of sci_name_qc.failureCacheMillis:
sci_name_qc.cacheSize=0stops results (including "no match" results) being cached, but failures are still remembered (unlesssci_name_qc.failureCacheMillis=0is also set), and threads looking up the same value at the same moment still share a single request.- The circuit breaker still stops calls to a service that keeps failing, whatever
sci_name_qc.failureCacheMillisis set to.
sci_name_qc.circuitBreakerThreshold=0disables it. An open circuit breaker can be closed withCircuitBreaker.forService(WoRMSService.SERVICE_NAME).reset()(orIRMNGService.SERVICE_NAME), orCircuitBreaker.resetAll(). - With both
sci_name_qc.failureCacheMillis=0andsci_name_qc.circuitBreakerThreshold=0, each lookup against a service that is down makes all of its attempts (by default four, with backoff between them) before failing.
Available in Maven Central.
<dependency>
<groupId>org.filteredpush</groupId>
<artifactId>sci_name_qc</artifactId>
<version>1.2.0</version>
</dependency>
To build an executable jar in the project directory (sci_name_qc-{version}-{commit}-executable.jar), run:
mvn package
To install in your local maven repository run:
mvn install
If javadoc errors in generated code are blocking a build and you want to be able to produce an artifact before addressing those, you can temporarily (these will still need to be addressed before deployment) suppress them to be just warnings with:
mvn clean package -DadditionalJOption=-Xdoclint:none
Test are separated into those be run offline and those that require conection to remote services (GBIF API, WoRMS aphia API).
Tests requiring online access to remote services are run in the integration-test phase, which lies between the package and install phases. So, invocation of
mvn package
will run the tests that do not need online access to services, (as will mvn test), while invocation of
mvn install
will run both these test and the integration tests that require access to online services (similarly mvn deploy).
It is currently safe to invoke mvn integration-test from the command line, but this is not a recommended practice,
as integration-tests are typically expected to start/stop a local jetty server in pre- and post- phases. If you
have to run mvn install in an offline environment, you can use mvn install -DskipTests to prevent test failures
from the absence of a connection to GBIF and WoRMS from causing the build to fail.
To deploy a snapshot to the snapshotRepository:
mvn clean deploy
To deploy a new release to maven central, set the version in pom.xml and in metadata to a non-snapshot version, then deploy with the release profile (which adds package signing and deployment to release staging:
-
set version in pom.xml
-
set version in Mechanism annotation in the DwCSciNameDQ classes in the generation configuration files, and in this README.
src/main/java/org/filteredpush/qc/sciname/DwCSciNameDQ.java src/main/java/org/filteredpush/qc/sciname/DwCSciNameDQDefaults.java generation/sci_name_qc_DwCSciNameQC_kurator_ffdq.config generation/sci_name_qc_DwCSciNameQC_stubs_kurator_ffdq.config README.md
-
deploy to maven central
mvn clean deploy -P release
After this, you may login to the sonatype oss repository hosting nexus instance find the staged release in the staging repositories and confirm the release.