ConjureRhyme sources and credits

ConjureRhyme is built on other people's data. This page says whose, what each part is used for, and under what terms. No source text is stored or shown: what the tool keeps are pronunciations and word counts.

Back to the search

Pronunciations

Every pronunciation comes from the Carnegie Mellon Pronouncing Dictionary (CMUdict), Copyright © 1993–2015 Carnegie Mellon University, maintained by the CMU Speech Group and distributed under BSD-style terms that ask users to acknowledge its origin. The tool uses a fork that adds a small number of pronunciations; errors in those are ours, not CMU's.

Sources

The “Source” setting chooses which body of English decides whether a phrase is one people actually use. Each source is stored as counts of words and of two- and three-word phrases; no sentence from any of them is kept.

Local debug builds of the tool also offer a fourth source, Peter Norvig's extract of the Google Web Trillion Word Corpus. The Linguistic Data Consortium distributes that corpus for research use, so this site does not serve it.

Word frequencies

A word that none of the sources records is weighed by wordfreq, by Robyn Speer, released under the Apache 2.0 licence with data under Creative Commons Attribution-ShareAlike 4.0. Its frequencies draw on Google Books Ngrams, OpenSubtitles, Wikipedia and the SUBTLEX word lists of Marc Brysbaert and colleagues, among other sources. Citation: Robyn Speer (2022), rspeer/wordfreq: v3.0, Zenodo, doi:10.5281/zenodo.7199437.

Word filters

The “Clean results only” filter starts from the List of Dirty, Naughty, Obscene, and Otherwise Bad Words, © Shutterstock, Inc., licensed under Creative Commons Attribution 4.0. The list of slurs removed from every search was assembled by hand from candidates proposed by Wiktionary's categories of slurs (Creative Commons Attribution-ShareAlike 3.0).

Rhythm

Which words may take or drop a beat freely follows the open- and closed-class split of Universal Dependencies, read off two annotated English treebanks: the English Web Treebank (Creative Commons Attribution-ShareAlike 4.0) and GUM, the Georgetown University Multilayer corpus (Creative Commons Attribution-NonCommercial-ShareAlike 4.0). The metrical rule they feed follows Paul Kiparsky, Stress, Meter, and Text-setting (2016).

Vowels

How alike two vowels count is derived from the formant measurements of J. Hillenbrand, L. A. Getty, M. J. Clark and K. Wheeler (1995), Acoustic characteristics of American English vowels, Journal of the Acoustical Society of America 97(5), 3099–3111, as hosted by Santiago Barreda with the authors' permission.