Automatic Detection of Spanish Neologisms
Automatic Detection of Spanish Neologisms
3, 382–399
doi: 10.1093/ijl/ecab009
Advance Access Publication Date: 20 June 2021
Article
Article
Rogelio Nazar
Pontificia Universidad Católica de Valparaı́so, Chile ([Link]@[Link])
Irene Renau
Pontificia Universidad Católica de Valparaı́so, Chile ([Link]@[Link])
Abstract
The appearance of new verbs can be observed regularly, but verbs are not frequent-
ly investigated in neology, and they are difficult to detect automatically. In this study,
a corpus-based method is proposed to detect Spanish verbs with a series of algo-
rithms that analyse the morphology of regular verbs. The vocabulary was drawn
from a large corpus and contrasted with a major dictionary of Spanish. Then, a ser-
ies of filters were applied to distinguish between valid neologism candidates and
spelling mistakes. Around 88% of the neologisms proposed by the method were cor-
rect and we estimate that the system detected 76% of the neologisms present in the
corpus. This procedure can be included in the workflow of a lexicographic project as
a regular part of the task, as a systematic way of collecting new verbs from the data
and avoiding under-representation or bias.
1. Introduction
In this paper, we present a methodological proposal for the detection of new Spanish verbs
in a general corpus, and for their incorporation in lexicographic tasks concerning the activ-
ity of updating dictionaries. Verbs are an open class of words and new units are regularly
added to the vocabulary, as we are observing now due to the coronavirus crisis: verbs such
as cuarentenar ‘to be in quarantine’ or desconfinar ‘to deconfine’ are new verbs created by
combining existing morphological structures in an innovative way. That is why Cook
(2010: 13) explains that ‘POS tagging of unknown words, including neologisms, can benefit
greatly from exploiting word structure’. However, morphological analysis is one of the
main sources of error in automatic neology detection, as many authors have reported.
acquired from other languages as loans (external neology). Furthermore, a new meaning
can be added to an already existing word. An example of this, called semantic neology,
would be the case of the verb rastrear ‘to follow the trace, to investigate’, which is currently
being used with the meaning of ‘tracking the contacts of a person who has been infected
with covid-19 in order to put them in quarantine’.
Different researchers have dealt with the theoretical and methodological considerations
As we have observed, dictionaries are both receivers of neologisms and material for
neology detection, working as stop lists. Concerning lexicographic projects, the fact that a
word is missing from the lemma list does not necessarily mean that it is a neologism, as the
criteria for the creation of the macrostructure of the dictionary involve many complex fac-
tors, including the type of user and their needs, the size and type of dictionary, the prescrip-
tive approach of the dictionary, the low frequency and dispersion of the word, etc. (Alvar
these units to be transitive in a higher percentage of those verbs included in the dictionary -
a tendency that, according to Bohrn (2010), can also be observed in Spanish verbs. Oliveira
(2020) deals with the process of creation of new verbs in Portuguese, which take place
mainly via suffixation, as it does in Spanish (and with equivalent suffixes: -ificar, -ear, -izar,
etc.).
L’Homme (2015) has argued that, in the case of terminology, verbs have been often
3. Methodology
The general approach of our method was to extract verbs from corpora and compare them
with the lemma list of a Spanish dictionary. The algorithm for the automatic extraction of
Spanish verbs from corpora is based on a set of rules to detect verbal inflection. The basic
idea was to hand-code a list of verb endings and to observe if the same root appeared in the
corpus in combination with multiple endings. In the case that it did, we reconstructed an in-
finitive form and contrasted it against a reference dictionary. If no match was found, we
submitted the candidate to a battery of tests to distinguish between true neologisms and
spelling mistakes.
3.1. Materials
Our materials comprised a corpus, for which we used the EsTenTen (Kilgarriff and Renau
2013) and a dictionary, used to contrast the verbs extracted from the corpus. As for the dic-
tionary, we used the Diccionario de la lengua espa~ nola, DLE (RAE 2014), mainly for tech-
nical reasons. The lemma list of this dictionary can be downloaded via Enclave RAE
([Link] a website that offers an advanced interface for the DLE. The list of
verbs from the dictionary comprised 11,893 verb lemmas.
The EsTenTen corpus consists of approximately 1010 running words from randomly
downloaded web pages of Spanish-speaking countries. This corpus is already tokenised, but
we ignored the tagging and used only the word forms. We only kept in the list those form
New Verbs and Dictionaries 387
types occurring at least 5 times in the corpus, a threshold we considered convenient to min-
imise accidental occurrences. We used 25% of this corpus to develop and test our method,
and the rest to evaluate its performance. From the first part we obtained 1,054,411 differ-
ent word forms, while from the rest of the corpus we obtained 2,954,973 different word
forms. It should be clear that what we mean by this is form types and not form tokens, and
also that form types are not to be confused with lemma types.
Enclitics
me te se nos le
les la lo las los
Table 2. Examples of verb root and endings from verbs rapear ‘to sing a rap’, retener ‘to retain’
and suprimir ‘to delete’.
Root Endings
Of course, what we needed to obtain was a list of verb lemmas and not a list of inflected
forms. Therefore, once this part of the procedure was finished, the next step was to find the
infinitive form of the verb, which in Spanish can only end in -ar, -er or -ir, as explained ear-
lier. The way to find the infinitive form was to attach the newly found root to each of these
endings and select the most frequent root-ending combination. For instance, in the case of a
verb root like reten-, shown in Table 2, the algorithm would test lemma candidates with
-ar: *retenar; with -er: retener; and with -ir: *retenir, but retener, the correct lemma, is the
one that appears most frequently in the corpus. Table 3 shows an example of the result of
this process, with the lemma, its frequency, root and endings. The DLE column states
whether the lemma is found in the dictionary.
If the newly found verb lemma matched the list obtained from the reference dictionary
(the DLE), then there was no doubt we were looking at a legitimate Spanish verb.
New Verbs and Dictionaries 389
Table 3. Examples of the reconstruction of the lemma forms of the verbs and their comparison
against the dictionary lemma list. Only rapear ‘to sing a rap’ was not present in the dictionary.
Prefixes. We used a list of prefixes in this process. If a prefix was detected in an infinitive,
and the verb was listed in the dictionary without the prefix, then the verb was considered a
neologism candidate, and then no other test was necessary. We adopted this strategy be-
cause one way for creating derivative neologisms is by adding prefixes to an existing word,
such as in recalendarizar ‘to re-calendarise’, from the existing verb calendarizar ‘calenda-
rise’ and the prefix re-. We extracted the list of prefixes from DLE, also by downloading it
from Enclave RAE. The full list is shown in Table 4.
Table 5, in turn, shows some examples of verbs not listed in the dictionary but bearing
one of the prefixes. As can be seen in the Frequency columns, the lemma without the prefix,
which does appear in the dictionary, is always far more frequent than the form with the
prefix.
Minimum length. Those candidates that did not match any prefix were subjected to the rest
of the battery tests. If a given candidate did not pass all these tests, it was discarded. The
first of such eliminating tests was to measure the length of the lemma in characters. As we
considered word endings of at least three letters, we discarded in this step all lemmas with
less than four characters. As explained by RAE and ASALE (2009), Spanish verbs tend not
to be too short because they carry at least four morphological segments: the root, the the-
matic vowel, the tense/mood and the person/number. Moreover, verbs with affixation and
enclitic pronouns are not infrequent, thus it is difficult to find verbs that can compress all
this morphological information in less than four letters. Examples of verb lemmas that
were eliminated in this step are *bar, *kir and *oir. None of these appear as verb lemmas
in the dictionary. The first form coincides with the noun bar ‘bar’, the second is not a word
and the last one resembles the Spanish verb oı́r ‘to hear’ but lacking the diacritical mark
due to misspelling.
390 Ana Castro et al.
ana-, anti-, auto-, cata-, cis-, co-, con-, contra-, cuasi-, de-, des-, di-, dia-, dis-, em-, en-, entre-, es-, ex-,
extra-, geo-, hiper-, hipo-, in-, infra-, inter-, intra-, para-, per-, peri-, pluri-, pos-, post-, pre-, pro-,
re-, requete-, res-, rete-, semi-, sin-, so-, sobre-, son-, sub-, super-, tele-, tera-, trans-, tras- ,ultra-
Reinsertion of first letter. Our second attempt to separate neologism candidates from spell-
ing mistakes was to reinsert a first letter, as we observed this is a very common occurrence.
Table 6 shows a few examples of false verb candidates that were obtained from the corpus,
the forms *aber, *econocer and *guantar. They certainly meet all the morphological criteria
described so far (e.g. they present a root which combines with different word endings) and
yet, they are not neologisms but spelling errors.
The idea was thus to try to reinsert an initial letter (such as a, e, i, o, u, h, b, c, d, f, g, j,
k, l, m, n, p, q, qu, r, s, t, v) and then check again against the reference dictionary. If a
match was found this time, then the candidate was discarded. Otherwise, it was submitted
to the following tests. For example, in the cases of Table 6, by adding an initial letter, the
system could match *aber, a mistake, with the real verb haber ‘to have’, which is not a
neologism candidate and was thus discarded.
Letter substitution. The remaining candidates were then screened for the detection of other
forms of typographic errors. A typical case of misspelling is the substitution of one letter
for another. For instance, a very common orthographic mistake in Spanish is the confusion
between b and v (e.g. it was common to find the verb avanzar ‘to move forward’ misspelled
as *abanzar). In order to detect and discard this type of error, we applied a special case of a
spell-checking algorithm consisting of a combination of letter substitution rules based on
Spanish orthography (RAE and ASALE 2010). Table 7 shows the entire catalogue of rules.
It should be read from left to right: for example, the letter or sequence of letters on the left
is replaced by the letter or sequence on the right.
If a letter substitution could be made in some verb candidate, the new form was then
compared with those in the dictionary. If a match was found, then the candidate was dis-
carded. Table 8 shows some examples of cases in which this happens.
Notice a very frequent case which is the incorrect addition of the letter h, a very com-
mon spelling mistake in Spanish because h is silent. The system eliminated this letter, result-
ing in the verb abrir ‘to open’, which is not a neologism. The table also shows the frequency
of the correctly spelled form in the corpus. If no match was found with these substitution
rules, the candidate was then submitted to the rest of the tests.
New Verbs and Dictionaries 391
Table 6. Examples of false candidates that were eliminated by inserting a first letter.
*aber h- haber
*econocer r- reconocer
*guantar a- aguantar
s!c n!m nb ! mb ps ! s n ! pn ó ! o
c!s s!z mb ! nb s ! ps pt ! t o ! ó
b!v z!s mb ! nv ii ! i t ! pt ú ! u
v!b z!c nv ! mb i ! ii bs ! s u ! ú
j!g y ! ll ep ! pe de ! des s ! bs u ! ü
g!j ll ! l pe ! ep di ! dis ns ! s ü ! u
c!k l ! ll p ! pe il ! ll s ! ns i!e
k!c ll ! y pe ! p li ! ll st ! s e!i
k ! qu r ! rr nf ! mf in ! ins s ! st a!e
qu ! k rr ! r x!j in ! inc á ! a e!a
n~ ! n h!- j!x oo ! o a ! á
n ! n~ y!i ee ! e o ! oo é ! e
N~ ! n~ i!y e ! ee gn ! n e ! é
ny ! n~ mp ! np x!s n ! gn ı́ ! i
m!n np ! mp s!x pn ! n i ! ı́
Table 8. Examples of misspelled verbs that were detected using substitution rules.
jA \ Bj
JðA; BÞ ¼ (1)
jA [ Bj
Again because of the computational cost that this comparison entails, we limited the ap-
plication of such coefficient only to those verbs in the dictionary that began with the same
letter as the candidate (except in the case of h, given that this is, as already explained, a
392 Ana Castro et al.
Table 9. Examples of verb candidates that matched and entry in the dictionary using an ortho-
graphic similarity coefficient.
very common occurrence), that were similar in length and which had a minimum frequency
of 15 occurrences in the corpus. Table 9 shows some examples of verb candidates that were
matched with verbs in the dictionary using the Jaccard similarity coefficient.
There we can see cases like *abadonar and *contatar, each of which found a match
with at least one entry in the dictionary using this coefficient: *abadonar matched abando-
nar ‘to abandon’, and contatar matched five possible verbs in the DLE. As in previous cases,
if at least one match with the dictionary was found, and such match was above the fre-
quency threshold, then the candidate was discarded. Otherwise, it continued to be sub-
jected to further tests.
To complement the previous similarity index, a new orthographic similarity coefficient
was implemented. As in the previous case, this new index takes two words as arguments
(the verb candidate and each of the entries in the dictionary) and proceeds to compare them
letter by letter. In a first loop, it compares the first letter of one with the first letter of the
other. If the letter was the same in both cases, a variable we call match was increased by
one, and then the algorithm moved to the next letter in both words. If, however, the letter
was not the same, then the first loop stopped and a second began, in a very similar fashion,
only that now instead of the first letter, it compared the last letter in both words. Again, if
the letter was the same, then the match variable was increased by one. Then, the pointer is
moved to the previous letter and the comparison is made again. The process finished when
another mismatch was found, now from the opposite direction. The result of the compari-
son was thus the quotient between the number of matches and the character length of the
larger word (2). As with the previous case, comparisons resulting in a value over an empir-
ically defined threshold were considered positive matches.
matchðA; BÞ
ortSimðA; BÞ ¼ (2)
max lengthðAÞ; lengthðBÞ
Table 10 shows two examples of verb candidates, *abodar and *colacar, that found a
match to some entries in the dictionary. The same criterion of frequency in the corpus was
used to avoid comparisons with very low frequency verbs of the dictionary.
and Vidal (2010). That is, if a word like flashear or bloguear shows a sharp rise in its fre-
quency of use, then this can be taken as a strong indication of neology. We leave this possi-
bility, however, for future work.
Table 12 presents the results of the overall evaluation, which informs how probable it is
that the correct true/false tag will be assigned to a given candidate.
We manually evaluated the total 4,349 verbs submitted to the battery of tests. We define
precision and recall in terms of true positives (tp), the case of a true neologism that is correctly
detected by the algorithm; false positives (fp), the case of a candidate that is not a neologism
but is anyway promoted by the algorithm; and finally false negatives (fn), the case of a neolo-
gism that was not detected by the algorithm (i.e., a true neologism that was discarded by
some filter). Using these three variables we calculated precision (3), which measures how like-
ly it is that a promoted candidate will be a true neologism, and recall (4), which measures the
exhaustivity of the method, i.e., how likely it is that a true neologism will be detected.
tp
precision ¼ (3)
ðtp þ fpÞ
tp
recall ¼ (4)
ðtp þ fnÞ
The overall precision of the detection of new verbs was 0.88 and recall was 0.76. Thus,
close to 88% of the final list of selected verbs are valid neologism candidates, resulting
from different word formation mechanisms.
With respect to the errors made by the algorithm, we found that they can be classified in
three groups: a) errors due to failure to detect missing accents (e.g. *sonreir was labelled a
neologism whereas it is a misspelling of sonreı́r ‘to smile’) and other typos (e.g. cosntruir
for construir ‘to build’); b) errors due to failure to detect verbs from other languages (e.g.
aproveitar, a Galician verb, was considered a valid Spanish verb), and c) errors due to in-
correct interpretation of a form as a verb, with the subsequent creation of a nonexistent
verb (e.g. *universar, *umbilicar).
Table 13 shows some examples of the list of neologism candidates, and the type of neol-
ogy mechanism according to Cabré (2015).
Some of these verbs are frequently used in everyday language, and the fact that they
have not been incorporated into the dictionary might be due to different factors, as
New Verbs and Dictionaries 395
Prefixation (adding a prefix to coorganizar, desactualizar, ‘to co-organise’, ‘to make some-
an existing verb) intercomunicar, precalcular, thing out-of-date’, ‘to inter-
reasignar, etc. communicate’, ‘to pre-calcu-
late’, ‘to reassign’, etc.
Suffixation (adding a verb suffix aperturar, chocolatear, coopera- ‘to open’, ‘to add chocolate’, ‘to
to an existing noun) tivizar, ficcionar, gelificar, etc. create an association’, ‘to
turn into fiction’, ‘to turn
into gel’, etc.
Parasynthesis (adding a prefix adinerar, desvirtualizar, enrutar, ‘to enrich’, ‘to devirtualise’, ‘to
and suffix to an existing etc. put in a route’, etc.
noun)
Loans (adding a Spanish verb customizar, draftear, loguear, ‘to customise’, ‘to make a
suffix to a loan) resetear, spoilear, etc. draft’, ‘to log’, ‘to reset’, ‘to
spoil’, etc.
discussed in section 2. Freixa and Torner (2020), among others, point out that there can be
a mismatch between the real institutionalisation of a word and the fact that the word is pre-
sent in the dictionary. If we observe the 100 most frequent candidates in our results, most
of them could be included in the dictionary, as they are frequent enough (they appear from
169 to 40,409 times in the corpus), they are morphologically correct and they are wide-
spread across Spanish-speaking countries. This is the case of vivenciar ‘to experiment’,
direccionar ‘to address’, googlear ‘to google’, suplementar ‘to supplement’, referenciar ‘to
refer’, resetear ‘to reset’, customizar ‘to customise’, etc.
It is not the goal of this study to try to find out why these verbs were not included in the
dictionary used to test the method. However, we can say that the resulting list of candi-
dates, with very few errors, would allow a lexicographic team to work with empirically-
based data and obtain valid information for potential new verb entries for dictionary
updating. Once the team had this information, it could make informed decisions about
which of these lemmas should be included and why, instead of using non-systematic ways
of including new words in dictionaries (e.g. following social trends or through personal
findings in newspapers). As already observed, there are many reasons why a new word
should or should not be included in a dictionary. Some of them may be based on the charac-
teristics of the dictionary project, combined with the characteristics of the potential new
lemma. For example, in the examples in Table 13, most of the verbs could be included in a
396 Ana Castro et al.
general descriptive dictionary, while aperturar, enrutar, customizar, draftear, loguear, rese-
tear or spoilear, being unnecessary loans, would not be adequate for a prescriptive
dictionary.
Finally, an interesting and unexpected result was the great number of verb entries from
the DLE that were not present in the large corpus we used in our experiments. From the
total of 11,893 verb lemmas of the dictionary, we found 6,208 that did not appear with the
5. Conclusions
In this paper, we proposed a method for the detection of new verbs, taking into account the
difficulty of part of speech tagging and the less-than-perfect performance of current taggers.
The proposed method, which is relatively simple, performs well in the detection of verbs in
Spanish, and differentiates verbs that are already present in a dictionary such as DLE from
the ones that the dictionary did not register. We also highlight the effectiveness of the
method at reducing the noise of spelling mistakes in the corpus. The typification of the
most frequent errors offers much help in reducing the burden of the otherwise manual revi-
sion of the data. Another characteristic of the proposed method is that it is based mainly on
morphological features to reduce the error rate, avoiding the use of semantic features,
which are harder to systematise.
The proposed algorithm can be easily incorporated into a workflow for creating or
updating a lexicographic database, so it can be used regularly to detect new verbs and con-
tribute to a more dynamic and updated dictionary. The vision of the electronic dictionary
as ‘up-to-date’ and a ‘dynamic repository of knowledge’ was already presented by De
Schryver (2003: 159), but almost 20 years later there is still plenty of work to do in this re-
spect. The automatic interrogation of a corpus to find verbs not present in a list of head-
words is a reliable method for tracing new verbs and adding them to a list of possible future
verb entries to the dictionary, aiding the lexicographer in her/his decision-making process.
Information such as the first appearance in the corpus, as well as the frequency and disper-
sion of the word, can also be included in the entry, in order to orient the user to the neologi-
cal nature of a specific unit. Depending on the type of dictionary and the conditions of the
lexicographic projects (such as technical, methodological, economical, etc.), a dictionary
for the Internet can be updated regularly with information about the frequency of the
word, taking data from corpora of different periods of times, comparing them and showing
the results in each entry. The general approach is not new (Renouf 1993), but it has not yet
been fully incorporated into the lexicographic procedures of Spanish dictionaries, nor is it a
standard procedure in other lexicographic traditions. As future work, we would like to add
supplementary information about the new verbs to the method. For example, the method
should be able to recognise whether a verb is only used in its participle form, or if it is used
with a specific alternation, among other features that could help lexicographers to add ac-
curate information to the dictionary entry and having first clues about its macrostructure.
New Verbs and Dictionaries 397
References
A. Dictionaries
Real Academia Espa~ nola. 2014. Diccionario de la lengua espa~ nola (Twenty-third edition.).
Madrid: Espasa.
B. Other literature
Abel, A. and E. Stemle. 2018. ‘On the Detection of Neologism Candidates as Basis for Language
Observation and Lexicographic Endeavours: The STyrLogism Project’ In Cibej, J., V. Gorjanc,
I. Kosem and S. Krek (eds), Proceedings of the XVIII EURALEX International Congress,
EURALEX 2018, Ljubljana, Slovenia, July 17-21, 2018. Ljubljana: University of Ljubljana,
535–544.
Adelstein, A. and I. Kuguel. 2008. De salariazo a corralito, de carapintada a blog. Nuevas palabras
en veinticinco a~nos de democracia. Los Polvorines: Universidad Nacional de General Sarmiento.
Algeo, J. 1993. ‘Desuetude among New English Words.’ International Journal of Lexicography
6.4: 281–293.
Alvar Ezquerra, M. 2007. ‘El neologismo espa~ nol actual’ In Luque Toro, L. (ed), Léxico espa~nol
actual. Actas del I Congreso Internacional de Léxico Espa~ nol Actual, Venice-Treviso, March
14-15, 2005. Venezia: Libreria Editrice Cafoscarina, 11–35.
Amore, M., S. McGregor and E. Jezek. 2018. ‘Distributional Analysis of Verbal Neologisms: Task
Definition and Dataset Construction’ In Cabrio, E., A. Mazzei and F. Tamburini (eds),
Proceedings of the Fifth Italian Conference on Computational Linguistics, CLiC-it 2018,
Torino, Italy. December 10-12, 2018. [Link], n. p.
Atkins, B. S., J. Kegl and B. Levin. 1988. ‘Anatomy of a Verb Entry: From Linguistic Theory to
Lexicographic Practice.’ International Journal of Lexicography 1.2: 84–126.
Atkins, B. S., and M. Rundell. 2008. The Oxford Guide to Practical Lexicography. Oxford:
Oxford University Press.
Baayen, R. H. and A. Renouf. 1996. ‘Chronicling the Times: Productive Lexical Innovations in an
English Newspaper.’ Language 72.1: 69–96.
Battaner, P. and S. Torner. 2008. ‘La polisemia verbal que muestra la lexicografı́a’ In Azorı́n D.,
B. Alvarado, J. Climent, M. I. Guardiola, R. Lavale Ortiz, C. Marimón Llorca, J. J. Joaquı́n
Martı́nez, X. A. Padilla, H. Provencio, I. Santamarı́a-Pérez, L. Timofeeva and E. Toro (eds).
Actas del II Congreso Internacional de Lexicografı́a Hispánica: el diccionario como puente
entre las lenguas y culturas del mundo. Alicante: Universidad de Alicante, 204–246.
Boas, F. 2001. ‘Frame Semantics as a Framework for Describing Polysemy and Syntactic
Structures of English and German Motion Verbs in Contrastive Computational Lexicography.’
In Rayson, P., A. Wilson, T. McEnery, A. Hardie and S. Khoja (eds), Proceedings of the Corpus
Linguistics 2001 Conference. Lancaster: Lancaster University, 64–73.
Bohrn, A. 2010. ‘La neologı́a verbal en el espa~ nol rioplatense.’ In Cabré, T., O. Domènech, R.
Estopà, J. Freixa and M. Lorente (eds), Actes del I Congrés Internacional de Neologia de les
Llengües Romàniques Barcelona: IULA, 917–927.
Cabré, M. T. 1999. Terminology. Theory, Methods and Applications. Amsterdam: John
Benjamins.
Cabré, M. T. and R. Estopà. 2009. ‘Trabajar en neologı́a con un entorno integrado en lı́nea: la
estación de trabajo OBNEO.’ Revista de Investigación Lingüı́stica 12: 17–38.
398 Ana Castro et al.
Cabré, M. T. and R. Nazar. 2012. ‘Towards a New Approach to the Study of Neology.’
Neologica 6: 63–80.
Cabré, M. T. 2015. ‘Bases para una teorı́a de los neologismos léxicos: primeras reflexiones’ In
Alves, I. and E. Sim~ oes (eds), Neologia das lı́nguas românicas. S~ ao Paulo: CAPES, Humanitas,
79–107.
Ca~nete, P., S. Fernández-Silva and B. Villena. 2019. ‘Estudio de los neologismos terminológicos
difundidos en el diario El Paı́s y su inclusión en el diccionario.’ Cı́rculo de Lingüı́stica Aplicada