Citation practices¶
Any publication, thesis, conference presentation, or preprint that uses ExELang data must follow these citation and disclosure guidelines. This protects participants, acknowledges contributors fairly, and maintains the project's data governance commitments.
Always cite the archives¶
All publications must cite the three scientific archives that make all of this work possible.
Homebank¶
All publications using a HomeBank corpus must include two citations: 1. the reference for the specific corpus used, as listed in the corpus documentation page on HomeBank. 2. the following reference for the archive itself: VanDam M, Warlaumont AS, Bergelson E, Cristia A, Soderstrom M, De Palma P., MacWhinney B. (2016). HomeBank: An online repository of daylong child-centered audio recordings. Seminars in Speech and Language, 37, 128–141. DOI: 10.1055/s-0036-1580745
For full citation rules, see: https://talkbank.org/homebank/rules.html
Databrary¶
Each dataset on Databrary has its own DOI and an automatically generated citation available on its volume page. The standard format is:
Author(s) (year). Title of dataset. Databrary. https://doi.org/10.17910/XXXXX
To find the citation for a specific dataset, navigate to its volume page on Databrary. For general information, see: https://databrary.org
The language Archive (TLA)¶
The language archive automatically generates a citation for each collection. The format is:
Author(s) (year). Title of collection. The Language Archive. https://hdl.handle.net/XXXXX (Accessed YYYY-MM-DD)
The citation is available directly on the collection page in the archive. For general information, see: https://archive.mpi.nl/tla
Model Development and Benchmarking Publications¶
When custodians publish work that falls within the defined scope of the consortium — that is, work specifically concerned with the development, training, evaluation, or benchmarking of automated models — they are entitled to do so without including corpus collectors as co-authors. The custodians' obligation in such cases is to: Cite each dataset used, following the citation guidelines outlined above for the relevant archive. Note that in rare cases, citing everyone may not be possible; see scenario 1 below
Example
Okko submits a paper at Interspeech presenting a neural network that has been trained using data from several corpora as well as several others. Each corpus is cited in the references. No corpus collector co-authorship is required because they made no contributions to the paper’s intellectual content.
Example
Situation: A custodian is submitting a paper to a conference which limits the number of citations to 10. It is therefore impossible for the custodian to comply with the requirement of citing all datasets and the three archives. Correct procedure: The custodian should cite the three archives in the main reference list and have an additional reference list including all dataset citations. The custodian should also email the corpus collectors so that they are aware of this unusual situation and don’t discover it later on.
Research Beyond Model Development: Consultation Requirement¶
When custodians intend to use consortium data for research that addresses substantive scientific questions beyond model performance — for example, cross-linguistic analysis of child-directed speech, or developmental trajectory modeling — they must notify each affected corpus collector in advance. The notification must:
- Describe the research question and the planned use of the data.
- Offer the corpus collector the option to withdraw their data from the specific project within a certain time frame. We suggest giving at least two weeks for the reply to arrive, and more over holiday times (which vary across countries, with some concentration around July-August in the Northern hemisphere, January-February in the Southern hemisphere).
- Offer the corpus collector the option to participate as a co-author, subject to their meeting standard co-authorship criteria as defined by the relevant journal or conference. Make sure to spell out what precise contributions are expected and the ideal timeline for them.
Example
Sho proposes a paper examining phonological variation in child-directed speech across multiple language communities using consortium corpora. Each corpus collector whose data is involved must be notified and offered the option to withdraw their data or join as co-author.
Bilateral Agreements¶
The rules in previous sections [a] and [b] represent the minimum baseline applicable to all ExELang Consortium members. Corpus collectors are free to negotiate separate bilateral agreements with custodians, modify, or extend these baseline terms.
Warning
This section is incomplete for the time being