<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Lexical Diversity &#8211; Alex Reuneker</title>
	<atom:link href="https://www.reuneker.nl/tag/lexical-diversity/feed/" rel="self" type="application/rss+xml" />
	<link>https://www.reuneker.nl</link>
	<description>Alex Reuneker&#039;s Blog</description>
	<lastBuildDate>Thu, 25 Dec 2025 07:39:20 +0000</lastBuildDate>
	<language>nl-NL</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.0.2</generator>
	<item>
		<title>&#167;290. Sorted frequency table for Hirsch-Popescu Point added to Lexical Diversity Calculator</title>
		<link>https://www.reuneker.nl/taal/sorted-frequency-table-for-hirsch-popescu-point-added-to-lexical-diversity-calculator/</link>
					<comments>https://www.reuneker.nl/taal/sorted-frequency-table-for-hirsch-popescu-point-added-to-lexical-diversity-calculator/#respond</comments>
		
		<dc:creator><![CDATA[Alex]]></dc:creator>
		<pubDate>Thu, 25 Dec 2025 07:39:20 +0000</pubDate>
				<category><![CDATA[Taal]]></category>
		<category><![CDATA[H-P Point]]></category>
		<category><![CDATA[Hirsch-Popescu Point]]></category>
		<category><![CDATA[Lexical Diversity]]></category>
		<category><![CDATA[Linguistics]]></category>
		<category><![CDATA[metric]]></category>
		<category><![CDATA[repetition]]></category>
		<guid isPermaLink="false">https://reuneker.nl/sorted-frequency-table-for-hirsch-popescu-point-added-to-lexical-diversity-calculator/?p=321</guid>

					<description><![CDATA[As I was working on a very brief piece on the Hirsch-Popescu Point (HPP) in one of the chapters of Reve&#039;s The Evenings, it occurred to me that]]></description>
										<content:encoded><![CDATA[<p>As I was working on a very brief piece on the <a href="https://www.reuneker.nl/2025/12/hirsch-popescu-point-added-to-lexical-diversity-calculator">Hirsch-Popescu Point</a> (HPP) in one of the chapters of Reve&#8217;s <a href="https://pushkinpress.com/book/the-evenings/" target="_blank" rel="noopener noreferrer">The Evenings</a>, it occurred to me that the <a href="https://www.reuneker.nl/files/ld">Lexical Diversity Calculator</a> does calculate the Hirsch-Popescu Point, but that it didn&#8217;t yet offer the option to actually look at the sorted frequency table used for determining the <a href="https://www.reuneker.nl/2025/12/hirsch-popescu-point-added-to-lexical-diversity-calculator">HPP</a>. As that table can be very informative, I now implemented the displaying of it in <a href="https://www.reuneker.nl/files/ld">Lexical Diversity Calculator</a>. The actual <a href="https://www.reuneker.nl/2025/12/hirsch-popescu-point-added-to-lexical-diversity-calculator">HPP</a>, so the word which has a position in the sorted frequency list that matches it frequency, is marked in bold and red for easy identification.</p>
<figure><img decoding="async" src="/wp-content/uploads/imports/20251225074311-Schermafbeelding%202025-12-25%20om%2007.30.21.png" alt="HPP marked in bold and red" /></figure>
<p><em>HPP marked in bold and red</em></p>
<p>If you&#8217;d like to use it, just head over to <a href="https://www.reuneker.nl/files/ld">https://www.reuneker.nl/files/ld</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://www.reuneker.nl/taal/sorted-frequency-table-for-hirsch-popescu-point-added-to-lexical-diversity-calculator/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>&#167;284. Hirsch-Popescu Point added to Lexical Diversity Calculator</title>
		<link>https://www.reuneker.nl/taal/hirsch-popescu-point-added-to-lexical-diversity-calculator/</link>
					<comments>https://www.reuneker.nl/taal/hirsch-popescu-point-added-to-lexical-diversity-calculator/#respond</comments>
		
		<dc:creator><![CDATA[Alex]]></dc:creator>
		<pubDate>Thu, 11 Dec 2025 12:28:35 +0000</pubDate>
				<category><![CDATA[Taal]]></category>
		<category><![CDATA[H-P Point]]></category>
		<category><![CDATA[Hirsch-Popescu Point]]></category>
		<category><![CDATA[Lexical Diversity]]></category>
		<category><![CDATA[Linguistics]]></category>
		<category><![CDATA[metric]]></category>
		<category><![CDATA[repetition]]></category>
		<guid isPermaLink="false">https://reuneker.nl/hirsch-popescu-point-added-to-lexical-diversity-calculator/?p=315</guid>

					<description><![CDATA[The Hirsch-Popescu Point (Popescu &#38;amp; Altmann, 2006) is an interesting metric to assess repetition in a text. It is determined by first]]></description>
										<content:encoded><![CDATA[<p>The <a href="https://www.geocities.ws/iipopescu/1_Some_aspects_of_word_frequencies.pdf" target="_blank" rel="noopener noreferrer">Hirsch-Popescu Point</a> (<a href="https://www.geocities.ws/iipopescu/1_Some_aspects_of_word_frequencies.pdf" target="_blank" rel="noopener noreferrer">Popescu &amp; Altmann, 2006</a>) is an interesting metric to assess repetition in a text. It is determined by first calculating the frequency distribution of all words in the text. Then, words are ranked from the most frequent to the least frequent. The H-P Point is then defined as &#8217;the point in which the ranking of a word in the distribution matches its frequency, just like the h-index in academia&#8217; (see <a href="https://dx.doi.org/10.2139/ssrn.2938838" target="_blank" rel="noopener noreferrer">Nunes, Ordanini, Valsesia, 2017</a>, p. 20; <a href="https://doi.org/10.1073/pnas.0507655102" target="_blank" rel="noopener noreferrer">Hirsh, 2005</a>). Indeed, the <a href="https://en.wikipedia.org/wiki/H-index" target="_blank" rel="noopener noreferrer">h-index</a> is a well-known measure of productivity and citation impact of publications. The smaller the H-P point, the less repetition a text contains and vice versa, i.e, the greater the HP-point, the more repetition a text contains.</p>
<p>If we apply the calculation to an example text, we easily see how it works exactly. The text used here is <a href="https://www.poetryfoundation.org/poems/49000/lady-lazarus" target="_blank" rel="noopener noreferrer">Sylvia Plath’s poem &#8216;Lady Lazarus&#8217;</a> (1965), and the resulting frequency table (distribution of words) can be seen below.</p>
<figure><img decoding="async" src="/wp-content/uploads/imports/20251211122951-Scherm%C2%ADafbeelding%202025-12-11%20om%2012.24.36.jpg" alt="Word distribution in https://www.poetryfoundation.org/poems/49000/lady-lazarus”&gt;Sylvia Plath’s poem ‘Lady Lazarus’ (1965)" /></figure>
<p><em>Word distribution in <a href="/%E2%80%9Chttps://www.poetryfoundation.org/poems/49000/lady-lazarus%E2%80%9D">Sylvia Plath’s poem ‘Lady Lazarus’</a> (1965)</em></p>
<p>In this table, we see that the word <em>of</em> occurs eight times in the poem, and it is also ranked at position eight in order from most to least frequent words. Therefore, 8 is the H-P Point. Of course, this frequency table, the sorting and determining the actual point at which frequency and order coincide is done by the <a href="https://www.reuneker.nl/files/ld">Lexical Diversity Calculator</a> for you, as can be seen below.</p>
<figure><img decoding="async" src="/wp-content/uploads/imports/20251211123105-Scherm%C2%ADafbeelding%202025-12-11%20om%2012.27.07.jpg" alt="The Hirsch-Popescu Point as calculated by the Lexical Diversity Calculator" /></figure>
<p><em>The <a href="https://www.geocities.ws/iipopescu/1_Some_aspects_of_word_frequencies.pdf" target="_blank" rel="noopener noreferrer">Hirsch-Popescu Point</a> as calculated by the <a href="https://www.reuneker.nl/files/ld">Lexical Diversity Calculator</a></em></p>
<p>As always, if you find it useful, have fun!</p>
]]></content:encoded>
					
					<wfw:commentRss>https://www.reuneker.nl/taal/hirsch-popescu-point-added-to-lexical-diversity-calculator/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>&#167;238. MATTR added to the Lexical Diversity Calculator</title>
		<link>https://www.reuneker.nl/taal/mattr-added-to-the-lexical-diversity-calculator/</link>
					<comments>https://www.reuneker.nl/taal/mattr-added-to-the-lexical-diversity-calculator/#respond</comments>
		
		<dc:creator><![CDATA[Alex]]></dc:creator>
		<pubDate>Wed, 04 Jun 2025 14:29:25 +0000</pubDate>
				<category><![CDATA[Taal]]></category>
		<category><![CDATA[calculator]]></category>
		<category><![CDATA[Lexical Diversity]]></category>
		<category><![CDATA[MATTR]]></category>
		<category><![CDATA[tokens]]></category>
		<category><![CDATA[ttr]]></category>
		<category><![CDATA[types]]></category>
		<guid isPermaLink="false">https://reuneker.nl/mattr-added-to-the-lexical-diversity-calculator/?p=287</guid>

					<description><![CDATA[Last week, I implemented the calculation of MATTR (Moving Average TTR) into the Lexical Diversity Calculator. MATTR calculates the mean TTR for]]></description>
										<content:encoded><![CDATA[<p>Last week, I implemented the calculation of <a href="https://www.sciencedirect.com/science/article/pii/S2772766124000740" target="_blank" rel="noopener noreferrer">MATTR (Moving Average TTR)</a> into the <a href="https://www.reuneker.nl/ld">Lexical Diversity Calculator</a>. MATTR calculates the mean TTR for successive windows of a text (<a href="https://www.tandfonline.com/doi/full/10.1080/09296171003643098" target="_blank" rel="noopener noreferrer">Covington &amp; McFall, 2010</a>), getting, at least that is the idea, a more stable indication of lexical diversity. While that’s not entirely the case (see <a href="https://www.sciencedirect.com/science/article/pii/S2772766124000740" target="_blank" rel="noopener noreferrer">Bestgen, 2025</a>), you can still test it at <a href="https://www.reuneker.nl/ld">https://www.reuneker.nl/ld</a>.</p>
<figure><img decoding="async" src="/wp-content/uploads/imports/20250604143422-sean-nufer-nupX62Vq7r0-unsplash.jpg" alt="enter image description here" /></figure>
<p><em>Photo by <a href="https://unsplash.com/@metalpsychologist?utm_content=creditCopyText&#038;utm_medium=referral&#038;utm_source=unsplash" target="_blank" rel="noopener noreferrer">Sean Nufer</a> on <a href="https://unsplash.com/photos/a-close-up-of-a-glass-window-with-the-words-slow-on-it-nupX62Vq7r0?utm_content=creditCopyText&#038;utm_medium=referral&#038;utm_source=unsplash" target="_blank" rel="noopener noreferrer">Unsplash</a></em></p>
<p>Next: implementing a compression-rate measure to operationalize text repetiveness for what hopefully becomes a project together with <a href="https://ivdnt.org/profile/vivien-waszink/" target="_blank" rel="noopener noreferrer">Vivien Waszink</a>!</p>
]]></content:encoded>
					
					<wfw:commentRss>https://www.reuneker.nl/taal/mattr-added-to-the-lexical-diversity-calculator/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>&#167;235. Improvements to the Lexical Diversity Calculator</title>
		<link>https://www.reuneker.nl/taal/improvements-to-the-lexical-diversity-calculator/</link>
					<comments>https://www.reuneker.nl/taal/improvements-to-the-lexical-diversity-calculator/#respond</comments>
		
		<dc:creator><![CDATA[Alex]]></dc:creator>
		<pubDate>Thu, 29 May 2025 08:58:42 +0000</pubDate>
				<category><![CDATA[Taal]]></category>
		<category><![CDATA[calculator]]></category>
		<category><![CDATA[Lexical Diversity]]></category>
		<category><![CDATA[MATTR]]></category>
		<guid isPermaLink="false">https://reuneker.nl/improvements-to-the-lexical-diversity-calculator/?p=285</guid>

					<description><![CDATA[In the last couple of days, I&#039;ve been implementing various improvements to the Lexical Diversity Calculator. Not only did I fix a problem in the]]></description>
										<content:encoded><![CDATA[<p>In the last couple of days, I&#8217;ve been implementing various improvements to the <a href="https://www.reuneker.nl/ld">Lexical Diversity Calculator</a>. Not only did I fix a problem in the calculation of <a href="https://link.springer.com/article/10.3758/BRM.42.2.381" target="_blank" rel="noopener noreferrer">MTLD</a>, which resulted in numbers that were slightly off, but I&#8217;ve also streamlined the calculations and added the calculation of <a href="https://www.sciencedirect.com/science/article/pii/S2772766124000740" target="_blank" rel="noopener noreferrer">Moving average TTR (MATTR)</a>.</p>
<figure><img decoding="async" src="/wp-content/uploads/imports/20250529090804-siora-photography-q5XDX1CadN4-unsplash.jpg" alt="enter image description here" /></figure>
<p><em>Photo by <a href="https://unsplash.com/@siora18?utm_content=creditCopyText&#038;utm_medium=referral&#038;utm_source=unsplash" target="_blank" rel="noopener noreferrer">Siora Photography</a> on <a href="https://unsplash.com/photos/black-text-q5XDX1CadN4?utm_content=creditCopyText&#038;utm_medium=referral&#038;utm_source=unsplash" target="_blank" rel="noopener noreferrer">Unsplash</a>.</em></p>
<p><strong>Updates</strong></p>
<ul>
<li>2025-05-29: Added choice to use natural logarithm or base 10 in</li>
</ul>
<p>calculation Maas&#8217;s a2, Dugast&#8217;s U2, and Herdan&#8217;s C.</p>
<ul>
<li>2025-05-29: Various improvements to calculations and algorithms; added <a href="https://www.sciencedirect.com/science/article/pii/S2772766124000740" target="_blank" rel="noopener noreferrer">MATTR</a>.</li>
<li>2025-05-26: Important change to the calculation of <a href="https://link.springer.com/article/10.3758/BRM.42.2.381" target="_blank" rel="noopener noreferrer">MTLD</a>, which was slightly off before due to not averaging the forward and backward algorithm.</li>
</ul>
<p>Next to this, I&#8217;m also working on an R-package to easily calculate several measures of lexical diversity, primarily for a research project I&#8217;m envisioning for the near future. Stay tuned! For now, please see the online calculator at <a href="https://www.reuneker.nl/ld">https://www.reuneker.nl/ld</a> for the newest version.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://www.reuneker.nl/taal/improvements-to-the-lexical-diversity-calculator/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>&#167;210. Hapax Legomena added to Lexical Diversity tool</title>
		<link>https://www.reuneker.nl/taal/hapax-legomena-added-to-lexical-diversity-tool/</link>
					<comments>https://www.reuneker.nl/taal/hapax-legomena-added-to-lexical-diversity-tool/#respond</comments>
		
		<dc:creator><![CDATA[Alex]]></dc:creator>
		<pubDate>Mon, 03 Mar 2025 12:21:16 +0000</pubDate>
				<category><![CDATA[Taal]]></category>
		<category><![CDATA[Hapax Legomena]]></category>
		<category><![CDATA[Lexical Diversity]]></category>
		<guid isPermaLink="false">https://reuneker.nl/hapax-legomena-added-to-lexical-diversity-tool/?p=271</guid>

					<description><![CDATA[In mailing back and forth with one of the researchers over at the [Max Planck Institute](https://www.mpi.nl/), there was some confusion over the use]]></description>
										<content:encoded><![CDATA[<p>In mailing back and forth with one of the researchers over at the <a href="https://www.mpi.nl/" target="_blank" rel="noopener noreferrer">Max Planck Institute</a>, there was some confusion over the use of the term <em>unique words</em> in the <a href="https://www.reuneker.nl/files/ld">Lexical Diversity too</a>l. Unique words are not <a href="https://en.wikipedia.org/wiki/Hapax_legomenon" target="_blank" rel="noopener noreferrer">hapax legomena</a>, which is the term in corpus linguistics for words that only occur once. Unique words are simply types and count up to the number of different words in a text. A word might occur once, twice or twenty times, but in all three cases, it would count as one unique word. This measure is also used for calculating the type-token-ratio. As the researcher was interested in how many words occur only once in a text, I&#8217;ve added this count. You can use the new feature <a href="https://www.reuneker.nl/files/ld">here</a> right away!</p>
<figure><img decoding="async" src="/wp-content/uploads/imports/20250303122630-Unknown.jpeg" alt="enter image description here" /></figure>
<p><em>Hapax legomena in the Lexical Diversity tool</em></p>
]]></content:encoded>
					
					<wfw:commentRss>https://www.reuneker.nl/taal/hapax-legomena-added-to-lexical-diversity-tool/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
	</channel>
</rss>
