<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>corpus &#8211; Alex Reuneker</title>
	<atom:link href="https://www.reuneker.nl/tag/corpus/feed/" rel="self" type="application/rss+xml" />
	<link>https://www.reuneker.nl</link>
	<description>Alex Reuneker&#039;s Blog</description>
	<lastBuildDate>Thu, 28 May 2026 14:12:25 +0000</lastBuildDate>
	<language>nl-NL</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.0.2</generator>
	<item>
		<title>&#167;347. Referentiecorpus SoNaR-500 toegevoegd aan Keyword Analysis-tool</title>
		<link>https://www.reuneker.nl/taal/referentiecorpus-sonar-500-toegevoegd-aan-keyword-analysis-tool/</link>
					<comments>https://www.reuneker.nl/taal/referentiecorpus-sonar-500-toegevoegd-aan-keyword-analysis-tool/#respond</comments>
		
		<dc:creator><![CDATA[Alex]]></dc:creator>
		<pubDate>Thu, 28 May 2026 14:12:25 +0000</pubDate>
				<category><![CDATA[Taal]]></category>
		<category><![CDATA[500]]></category>
		<category><![CDATA[analysis]]></category>
		<category><![CDATA[corpus]]></category>
		<category><![CDATA[keyword]]></category>
		<category><![CDATA[referentie]]></category>
		<category><![CDATA[sonar]]></category>
		<guid isPermaLink="false">https://reuneker.nl/referentiecorpus-sonar-500-toegevoegd-aan-keyword-analysis-tool/?p=353</guid>

					<description><![CDATA[Gisteren en vandaag was ik bezig met een woordfrequentielijst van het SoNaR-corpus (Oostdijk et al., 2013). Die lijst heb ik nodig om de Lexical]]></description>
										<content:encoded><![CDATA[<p>Gisteren en vandaag was ik bezig met een woordfrequentielijst van het <a href="https://taalmaterialen.ivdnt.org/download/tstc-sonar-corpus/" target="_blank" rel="noopener noreferrer">SoNaR-corpus</a> (<a href="https://research.tilburguniversity.edu/en/publications/sonar-500-2/" target="_blank" rel="noopener noreferrer">Oostdijk et al., 2013</a>). Die lijst heb ik nodig om de <a href="https://www.reuneker.nl/files/ld/">Lexical Diversity-tool</a> uit te breiden, maar ik heb het <a href="https://taalmaterialen.ivdnt.org/download/tstc-sonar-corpus/" target="_blank" rel="noopener noreferrer">SoNaR</a> vast als referentiecorpus toegevoegd aan de <a href="https://www.reuneker.nl/files/keyword/">Keyword Analysis-tool</a>.</p>
<p>Je kunt nu dus kiezen om trefwoorden in je (Nederlandse) tekst op te sporen door de tekst te vergelijken met het toch wel oude en veel kleinere <a href="https://www.dbnl.org/tekst/_ned020200001_01/_ned020200001_01_0033.php" target="_blank" rel="noopener noreferrer">CONDIV-corpus</a>, of met het <a href="https://taalmaterialen.ivdnt.org/download/tstc-sonar-corpus/" target="_blank" rel="noopener noreferrer">SoNaR-corpus</a>. Andere beschikbare referentiecorpora zijn het BNC voor het (Brits) Engels, een Nederlandstalig popcorpus en een eveneens Nederlandstalig rapcorpus. Je kunt uiteraard ook nog steeds zelf een referentiecorpus toevoegen &#8212; dat is makkelijker dan je wellicht denkt!</p>
<p>In de onderstaande afbeelding kun je zien dat bijvoorbeeld het woord <em>herkomstlanden</em> significant vaker voorkomt in het NOS-artikel <a href="https://nos.nl/artikel/2616154-onderzoek-deel-collectie-oranjes-mogelijk-onrechtmatig-verkregen" target="_blank" rel="noopener noreferrer">Onderzoek: deel collectie Oranjes mogelijk onrechtmatig verkregen</a> dan in het SoNaR-corpus en dus iets zegt over de het artikel; het is een trefwoord of <em>keyword</em>.</p>
<figure><img decoding="async" src="/wp-content/uploads/imports/20260528144727-Scherm%C2%ADafbeelding%202026-05-28%20om%2014.47.03.jpg" alt="Trefwoorden in vergelijking met het SoNaR-corpus" /></figure>
<p><em>Trefwoorden in vergelijking met het SoNaR-corpus</em></p>
<p>Opmerkingen bij deze toevoeging zijn dat alleen Nederlandse krantenteksten zijn gebruikt voor de frequentielijst en, met het oog op <em>processing</em> in <em>JavaScript</em> en bestandsgroottes, alleen woorden die tien keer of vaker voorkwamen zijn meegenomen.</p>
<p>Je kunt de uitgebreide tool uiteraard direct gebruiken op <a href="https://www.reuneker.nl/files/keyword/">https://www.reuneker.nl/files/keyword</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://www.reuneker.nl/taal/referentiecorpus-sonar-500-toegevoegd-aan-keyword-analysis-tool/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>&#167;296. Dealing with &#039;zero counts&#039; in keyword analysis</title>
		<link>https://www.reuneker.nl/taal/dealing-with-zero-counts-in-keyword-analysis/</link>
					<comments>https://www.reuneker.nl/taal/dealing-with-zero-counts-in-keyword-analysis/#respond</comments>
		
		<dc:creator><![CDATA[Alex]]></dc:creator>
		<pubDate>Tue, 30 Dec 2025 06:06:55 +0000</pubDate>
				<category><![CDATA[Taal]]></category>
		<category><![CDATA[corpus]]></category>
		<category><![CDATA[keyword analysis]]></category>
		<category><![CDATA[reference]]></category>
		<category><![CDATA[target]]></category>
		<category><![CDATA[zero counts]]></category>
		<guid isPermaLink="false">https://reuneker.nl/dealing-with-zero-counts-in-keyword-analysis/?p=326</guid>

					<description><![CDATA[One problem with keyword analysis is that the target corpus will likely include words that do not occur in the reference corpus. In calculating]]></description>
										<content:encoded><![CDATA[<p>One problem with keyword analysis is that the target corpus will likely include words that do not occur in the reference corpus. In calculating various measures of <em>keyness</em>, this would result in a division by zero, which is mathematically impossible, as far as I know. The default way of dealing with this is to assign words that do not occur in the reference corpus a frequency of 0.5, but this introduces the risk of a result in which  such keywords dominate the top positions, because their <em>keyness</em> is inflated.</p>
<p>To remedy this problem, I have added an option to the <a href="https://www.reuneker.nl/files/keyword/">Keyword Analysis Tool</a> which let&#8217;s you choose to either go with the default of assigning a 0.5 frequency to &#8216;zero counts&#8217;, or to simply discard them from all calculations, resulting in keywords that have a minimal frequency of 1 in the reference corpus.</p>
<figure><img decoding="async" src="/wp-content/uploads/imports/20251230061633-Schermafbeelding%202025-12-30%20om%2006.16.19.png" alt="Dealing with &#x27;zero counts&#x27; in the Keyword Analysis Tool" /></figure>
<p><em>Dealing with &#8216;zero counts&#8217; in the <a href="https://www.reuneker.nl/files/keyword/">Keyword Analysis Tool</a></em></p>
<p>There is no real wrong or right way to do this, but at least now you have a choice. Have fun!</p>
]]></content:encoded>
					
					<wfw:commentRss>https://www.reuneker.nl/taal/dealing-with-zero-counts-in-keyword-analysis/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>&#167;225. &#039;From clue to culprit: epistemic conditionals in detective fiction&#039; at ECA 2025, Warsaw</title>
		<link>https://www.reuneker.nl/taal/from-clue-to-culprit-epistemic-conditionals-in-detective-fiction-at-eca-2025-warsaw/</link>
					<comments>https://www.reuneker.nl/taal/from-clue-to-culprit-epistemic-conditionals-in-detective-fiction-at-eca-2025-warsaw/#respond</comments>
		
		<dc:creator><![CDATA[Alex]]></dc:creator>
		<pubDate>Tue, 29 Apr 2025 07:33:31 +0000</pubDate>
				<category><![CDATA[Taal]]></category>
		<category><![CDATA[argumentation]]></category>
		<category><![CDATA[conditionals]]></category>
		<category><![CDATA[corpus]]></category>
		<category><![CDATA[detectives]]></category>
		<category><![CDATA[reasoning]]></category>
		<category><![CDATA[sherlock holmes]]></category>
		<guid isPermaLink="false">https://reuneker.nl/from-clue-to-culprit-epistemic-conditionals-in-detective-fiction-at-eca-2025-warsaw/?p=281</guid>

					<description><![CDATA[In September 2025, at the 5th European Conference on Argumentation (ECA), I will present a corpus study on reasoning with conditionals by detectives,]]></description>
										<content:encoded><![CDATA[<p>In September 2025, at the <a href="https://ecargument.org/" target="_blank" rel="noopener noreferrer">5th European Conference on Argumentation (ECA)</a>, I will present a corpus study on reasoning with conditionals by detectives, entitled <em>From clue to culprit: epistemic conditionals in detective fiction</em>. The conference is hosted by the <a href="https://eng.ans.pw.edu.pl/" target="_blank" rel="noopener noreferrer">Warsaw University of Technology, Faculty of Administration and Social Sciences</a>. Below you&#8217;ll find the abstract. Hope to see you all there!</p>
<figure><img decoding="async" src="/wp-content/uploads/imports/20250429073644-Scherm%C2%ADafbeelding%202025-04-29%20om%2007.36.35.jpg" alt="enter image description here" /></figure>
<p><em>Abstract &#8216;From clue to culprit: epistemic conditionals in detective fiction&#8217;</em></p>
]]></content:encoded>
					
					<wfw:commentRss>https://www.reuneker.nl/taal/from-clue-to-culprit-epistemic-conditionals-in-detective-fiction-at-eca-2025-warsaw/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>&#167;215. Random Text Sampler</title>
		<link>https://www.reuneker.nl/taal/random-text-sampler/</link>
					<comments>https://www.reuneker.nl/taal/random-text-sampler/#respond</comments>
		
		<dc:creator><![CDATA[Alex]]></dc:creator>
		<pubDate>Thu, 20 Mar 2025 17:51:53 +0000</pubDate>
				<category><![CDATA[Taal]]></category>
		<category><![CDATA[corpus]]></category>
		<category><![CDATA[online]]></category>
		<category><![CDATA[random]]></category>
		<category><![CDATA[sample]]></category>
		<category><![CDATA[text]]></category>
		<category><![CDATA[tool]]></category>
		<guid isPermaLink="false">https://reuneker.nl/random-text-sampler/?p=275</guid>

					<description><![CDATA[Soms is het handig om voor een vergelijkend onderzoek steekproeven (*samples*) van een bepaald aantal woorden uit een tekst te halen. Omdat dat]]></description>
										<content:encoded><![CDATA[<p>Soms is het handig om voor een vergelijkend onderzoek steekproeven (<em>samples</em>) van een bepaald aantal woorden uit een tekst te halen. Omdat dat typisch zo’n terugkerend klusje is waaraan ik elke keer toch weer meer tijd kwijt ben dan gedacht, heb ik er maar een <em>online tooltje</em> voor gemaakt.</p>
<figure><img decoding="async" src="/wp-content/uploads/imports/20250323141624-Scherm%C2%ADafbeelding%202025-03-23%20om%2014.16.07.jpg" alt="enter image description here" /></figure>
<p><em>Random text sampler</em></p>
<p>Het lijkt me zonde om dat voor mezelf te houden en daarom kan iedereen die dat wil op <a href="https://www.reuneker.nl/randsamples">https://www.reuneker.nl/randsamples</a> een tekst invoeren, het gewenste aantal steekproeven en de steekproefgrootte (in aantal woorden) selecteren en met een druk op de knop de <em>samples</em> tevoorschijn toveren. Je kunt daarbij ook aangeven dat je, per <em>sample</em> en voor het geheel, de <a href="https://en.wikipedia.org/wiki/Lexical_diversity" target="_blank" rel="noopener noreferrer">*type-token-ratio’s*</a> en <a href="https://link.springer.com/article/10.3758/BRM.42.2.381" target="_blank" rel="noopener noreferrer">MTLD</a>-scores wilt zien.</p>
<p>Concreet was de aanleiding overigens een klein onderzoekje naar jeugdliteratuur ter illustratie van de <a href="https://www.reuneker.nl/t">t-toets-calculator</a> voor studenten, dat je hier vindt: <a href="https://www.reuneker.nl/files/blog/2025/03/zinslengte-in-de-brief-voor-de-koning-en-kinderen-van-moeder-aarde">https://www.reuneker.nl/files/blog/2025/03/zinslengte-in-de-brief-voor-de-koning-en-kinderen-van-moeder-aarde</a>. Mocht je gewoon eens willen kijken hoe e.e.a. werkt, dan kun je gemakkelijk samples nemen uit <a href="https://www.gutenberg.org/ebooks/164" target="_blank" rel="noopener noreferrer">Jules Vernes *Twenty Thousand Leagues under the Sea*</a> of <a href="https://gutenberg.org/ebooks/67219" target="_blank" rel="noopener noreferrer">Louis Couperus&#8217;  TOK0 Stille Kracht TOK1 </a>, die je met een klik op de desbetreffende knop op het scherm tovert.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://www.reuneker.nl/taal/random-text-sampler/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>&#167;165. New publication in Argumentation: Assessing Classification Reliability&#8230;</title>
		<link>https://www.reuneker.nl/taal/new-publication-in-argumentation-assessing-classification-reliability-of-conditionals-in-discourse/</link>
					<comments>https://www.reuneker.nl/taal/new-publication-in-argumentation-assessing-classification-reliability-of-conditionals-in-discourse/#respond</comments>
		
		<dc:creator><![CDATA[Alex]]></dc:creator>
		<pubDate>Fri, 07 Apr 2023 13:32:00 +0000</pubDate>
				<category><![CDATA[Taal]]></category>
		<category><![CDATA[argumentation]]></category>
		<category><![CDATA[classification]]></category>
		<category><![CDATA[conditionals]]></category>
		<category><![CDATA[corpus]]></category>
		<category><![CDATA[inter-rater reliability]]></category>
		<category><![CDATA[Linguistics]]></category>
		<category><![CDATA[research]]></category>
		<guid isPermaLink="false">https://reuneker.nl/new-publication-in-argumentation-assessing-classification-reliability-of-conditionals-in-discourse/?p=251</guid>

					<description><![CDATA[Different types and argumentative uses of conditionals (if-then) have been distinguished in the literature, but their applicability to actual]]></description>
										<content:encoded><![CDATA[<p>Different types and argumentative uses of conditionals (if-then) have been distinguished in the literature, but their applicability to actual language use is rarely evaluated.</p>
<p>As &#8217;the proof of the pudding is in the eating&#8217;, my new paper in Argumentation (Springer) entitled &#8216;Assessing Classification Reliability of Conditionals in Discourse&#8217; addresses this issue by means of an experiment in which the inter-rater reliability of classifications applied to natural-language corpora was assessed.</p>
<figure><img decoding="async" src="/wp-content/uploads/imports/20230407133136-Scherm%C2%ADafbeelding%202023-04-07%20om%2013.22.20.jpg" alt="enter image description here" /></figure>
<p><em>New publication in Argumentation: &#8216;Assessing Classification Reliability of Conditionals in Discourse&#8217;</em></p>
<p>You can find the paper (open access) in Argumentation here: <a href="https://rdcu.be/c9nO4" target="_blank" rel="noopener noreferrer">https://rdcu.be/c9nO4</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://www.reuneker.nl/taal/new-publication-in-argumentation-assessing-classification-reliability-of-conditionals-in-discourse/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>&#167;112. Custom reference corpora in keyword analysis</title>
		<link>https://www.reuneker.nl/taal/custom-reference-corpora-in-keyword-analysis/</link>
					<comments>https://www.reuneker.nl/taal/custom-reference-corpora-in-keyword-analysis/#respond</comments>
		
		<dc:creator><![CDATA[Alex]]></dc:creator>
		<pubDate>Wed, 23 Nov 2022 15:22:52 +0000</pubDate>
				<category><![CDATA[Taal]]></category>
		<category><![CDATA[analysis]]></category>
		<category><![CDATA[corpus]]></category>
		<category><![CDATA[custom]]></category>
		<category><![CDATA[keyword]]></category>
		<guid isPermaLink="false">https://reuneker.nl/custom-reference-corpora-in-keyword-analysis/?p=248</guid>

					<description><![CDATA[Today I added the option to directly compare two texts on the [keyword analysis page][1]. Before today, only one general Dutch and one general]]></description>
										<content:encoded><![CDATA[<p>Today I added the option to directly compare two texts on the <a href="https://www.reuneker.nl/files/keyword/">keyword analysis page</a>.</p>
<p>Before today, only one general Dutch and one general English reference corpus could be loaded, but much of the time, a custom corpus is needed to get more informative results. For example, say you&#8217;d like to see a list of keywords in a certain novel. It makes sense to compare this novel to another novel, as in the screenshot below, or perhaps to a collection of other novels.</p>
<figure><img decoding="async" src="/wp-content/uploads/imports/20221123152246-Scherm%C2%ADafbeelding%202022-11-23%20om%2015.22.21.png" alt="enter image description here" /></figure>
<p>Well, now you can. Simply copy-paste the reference corpus to the webpage, and you&#8217;re good to go.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://www.reuneker.nl/taal/custom-reference-corpora-in-keyword-analysis/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
	</channel>
</rss>
