Post

HN
Hacker News

English: A vs. An

Ranked #3 on Hacker News with 139 points and 177 comments.

In English, there is an “indefinite” article a that can go before a word. For example, a raccoon . But for some words, we use an . For example, an apple .

When procedurally generating text, I want a function a_or_an("apple") that tells me which article to use. That seems like it’d be easy. We can check the first letter to see if it’s a vowel. But that would mean we output an unicorn , not a unicorn .

The actual rule is not whether the written word starts with a vowel letter, but whether the spoken word starts with a vowel sound. The word <unicorn> starts with vowel letter ( <u> ) but a consonant sound (cmudict Y , ipa /j/ ). The word <hour> starts with a consonant letter ( <h> ) but a vowel sound (cmudict AW , ipa /aʊ/ ).

I was curious how often these exceptions occurred, and whether they can be grouped together, so I spent a day looking at the data and building some visualizations and wrote up the results. I was surprised that only 129 of the 32,455 words in my list needed exceptions.

[LLM note: I did not use LLMs to write any of this code, but in hindsight, I should have. This is one-off code to answer a question. It doesn’t need to be clean or maintainable. It only needs to be correct. I would’ve spent more time on the trie simplification algorithm and less time on parsing cmudict and re-learning d3.js.]