Multi-Lingual Question Generation Model
Multi-Lingual Question Generation Model
Bingning Wang, Ting Yao, Weipeng Chen, Jingfang Xu and Xiaochuan Wang
Sogou Inc.
Beijing, China
{wangbingning,yaoting}@[Link]
2262
Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 2262–2272
August 1–6, 2021. ©2021 Association for Computational Linguistics
tain some unnecessary language-specific features role of the passage (Mostow and Chen, 2009; Heil-
that are too specific to transfer across languages. man and Smith, 2010). First, the target answers are
Inspired by previous works on transfer learning generated using rules or semantic roles, next, low-
(Chen et al., 2017; Liu et al., 2017), we propose an quality questions are generated using hand-crafted
adversarial training objective to decouple the low- rules or templates. Finally, the generated questions
level module with the high-level module, which are ranked by features such as keyword matching
prevents the private and shared latent feature spaces degree or sentence perplexity (Hussein et al., 2014).
from interfering with each other, making the high- The main drawbacks of these symbolic systems are
level module language-invariant, thus achieving that the rules and templates are expensive to manu-
better transferability for different languages. ally create, and lack diversity.
To get a better initialization for our model, we de-
With the development of deep learning and large-
velop two self-supervised methods to pre-train our
scale question answering datasets, motivated by
model on abundant monolingual text. We apply
neural machine translation, Du et al. (2017) pro-
our model to five languages QG tasks that have
posed a sequence to sequence (seq2seq) architec-
human-labeled QG datasets. The experimental
ture combined with attention mechanism, achiev-
results demonstrate that all languages QG could
ing a promising result on QA dataset SQuAD.
benefit from the multi-lingual training. Our mod-
Since then, many works have been proposed to
els surpass previous monolingual or multi-lingual
extend the preliminary framework with rich fea-
QG methods by a large margin, even in zero-shot
tures, such as named entity tags (Zhou et al., 2017)
learning where we had no training data in the low-
or answer position features (Duan et al., 2017), and
resource languages, our model achieves satisfac-
incorporate copy mechanism to copy words from
tory results by merely trained on English dataset,
the context paragraph (Song et al., 2018). Other
which shows a promising transferability of the pro-
types of models are also introduced such as graph
posed model.
neural networks (Chen et al., 2019) or Transformer
Besides, we also propose a large-scale Chinese
(Scialom et al., 2019). However, most of these
QG dataset containing more than 220k human-
works are focus on English QG and have not been
labeled questions. We hope the proposed Chi-
validated in other languages.
nese dataset could benefit the community for more
comprehensive multi-lingual QG research. The Multi-Lingual language generation. Duan
codes and proposed datasets are available at https: et al. (2019) translated documents as weakly su-
//[Link]/benywon/LALM. pervised training data for zero-shot multi-lingual
Our contributions are summarized as follow: abstractive summarization. Chi et al. (2019) pro-
posed a multi-lingual pre-training method that can
• We propose a novel language-agnostic language transfer monolingual supervision signals to other
model which decouples the language specific pre-trained languages. Zhu et al. (2019) adopt
and language independent information in QG. large-scale supervised data from existing monolin-
• The proposed model achieves significant im- gual summarization datasets via translation strategy
provement over previous models in multi-lingual to perform multi-lingual summarization. Kumar
QG, and we analyze the transferability in multi- et al. (2019) also proposed a multi-lingual question
ple languages. generation methods based on Transformer, they
• We release a large-scale human labeled Chinese proposed a small Hindi QG dataset and improved
QG dataset containing more than 220k questions. the QG result on Hindi by training with additional
To our best knowledge, this is the largest specific English data.
question generation dataset so far.
Compared with the previous multi-lingual meth-
2 Related Work ods, our method directly separates the language-
dependent module and language-independent mod-
Question generation has received increasing at- ule. We propose an adversarial decoupling module
tention from the research community. Traditional to improve the adaptive ability of the model. Be-
QG systems are mostly rule-based, which some- sides, our model could be properly pre-trained by
times utilizing off-the-shelf tools to get the syntac- monolingual data, which obviates the need to con-
tic structure, dependency relations, and semantic struct the back-translation or pseudo-parallel data.
2263
English
Hindi Low-Level
Linear+Softmax Chinese En 0.3
×L
Shared Zh 0.4
High-Level
Block maximize
Fr 0.2
LayerNorm
Not mask
…
Linear Masked
Embedding+LSTM
Embedding+LSTM
x1, . . . , x|x|, sep, y1, . . . , y|y|
En
Figure 1: (a) The whole architecture of the proposed language-agnostic language model. It consists of the low-
level language understanding module (Embedding+LSTM) and the high-level semantic understanding module
(Transformer block), followed by a projection and softmax module. (b) The attention mask matrix M in the high-
level module, Mi,j means whether the word in position i could attend to the word in position j. The gray cells
are allowed to attend and the others are masked to forbid attention. (c) The adversarial decoupling module where
the discriminator tries to maximize the probability of the corresponding language while the generator (low-level
module) tries to minimize it.
3 Language-agnostic Language Model separating the low-level module for each language
could benefit a lot for multi-lingual QG.
The Language-agnostic language model (LALM)
consists of the low-level module and the high-level 3.2 High-Level Module
module, the whole architecture is illustrated in Fig-
The low-level module is built to perform the basic
ure 1(a) we describe it below.
linguistic understanding, and the high-level mod-
3.1 Low-Level Module ule is built on top of the low-level module to per-
form higher-level information aggregation, which
The low-level module is built to perform the basic
requires higher model capacity. In this paper, we
language understanding. In this paper, we adopt
use the Transformer (Vaswani et al., 2017) model
the LSTM (Hochreiter and Schmidhuber, 1997) en-
as the high-level module.
coder as the low-level language understanding mod-
The Transformer, with the core building-block
ule1 . LSTM processes text in sequential order and
called multi-head attention, has shown great ad-
embeds the language information into dense repre-
vantages in representing languages in many NLP
sentations. We adopt the uni-directional LSTM in
tasks. Current state-of-the-art models in natural
this paper to make the model auto-regressive.
language understanding benchmark GLEU2 (Wang
For the language-agnostic language models,
et al., 2018) are almost Transformer-based. In this
each language has its specific word embeddings
paper, we focus on QG which is a sequence-to-
and specific low-level language understanding
sequence problem, so we adopt the mask operation
LSTMs. This is different from some previous
similar with (Dong et al., 2019), which is illustrated
multi-lingual methods that a shared or aligned word
in Figure 1(b). For a pair of sequence (x, y) where
embedding is utilized for different languages (Con-
x = x1 , ..., x|x| is the source, and y = y1 , ..., y|y|
neau et al., 2018; Lample and Conneau, 2019). Sep-
is the target, we concatenate them together with a
arating the language understanding module enables
special token <sep>, forming a single sequence
us to model specific linguistic characteristics in dif-
with length |x| + |y| + 1. We want all the positions
ferent languages. In Section 4, we will show that
in the source {1, 2, ..., |x|} to attend to each other
1
In fact, we also conduct experiments on adopting other so we can obtain the bi-directional representations
types of models as the low-level module such as Transformer
2
or GRU, but the result is not comparable with the LSTM. [Link]
2264
of the source, and all the positions in the target representations without discriminative language in-
{|x| + 1, ..., |x| + |y| + 1} are forbidden to attend formation. In this way, the discriminator acts as
to future words: an adversarial decoupling module (ADM) to en-
( courage the low-level module to generate language-
0, j ≤ |x| or j ≤ i, agnostic representations.
Mi,j = (1)
− ∞, otherwise. The architecture of ADM is shown in Figure 1
(c), and the loss function for the discriminator and
This attention mask operation enables us to build low-level module (generator) are:
a causal language model that the generation of the
current word only depends on its previous words. LD = − log p(ŷi )
Therefore, the probability of y could be denoted as: (5)
LG = − log p(1 − ŷi )
|y|
Y where ŷi is the discriminator probability for the in-
p(y|x) = p(yi |y<i , x) (2) put language i. In fact, the objective of the genera-
i=1
tor is to maximize the entropy of the discriminator’s
And the loss for the whole model is the negative output to make it less confident of the language.
log likelihood of the data:
X 3.4 Pre-training
LN LL = − E log p(yi |y<i , x) (3) Recent works on NLP and language generation
x,y
have shown the great advantage of large-scale
3.3 Adversarial Decoupling Module pre-training (Devlin et al., 2018; Radford et al.,
In this paper, we want the representations of the 2019; Lewis et al., 2019; Roberts et al., 2020). In
low-level module in different languages to con- this paper, we also pre-train our model in mas-
tain no language-specific information that is inter- sive multilingual text. Since our model is a se-
leaved with the high-level module. In this way, quence to sequence architecture, we develop two
the high-level module could focus on the semantic self-supervised objectives for language generation
understanding shared across languages. We build pre-training:
a discriminator on top of the low-level module to Denoised Auto-Encoder (DAE): Most previous
determine whether the output of low-level represen- works on natural language generation pre-training
tations contains the specific language information. resort to DAE to initialize the model. In DAE, a
The discriminator is a bi-directional LSTM tak- corrupted version of the original sentence is cre-
ing the output of the low-level module as input and ated as the source and the model should reconstruct
tries to predict its language. Concretely, denote the the original sentence. In this paper, we adopt the
output of the low-level module is S ∈ Rn,d where similar noising strategy as Lewis et al. (2019): (1)
n is the sequence length (i.e. |x| + |y| + 1), and Token Masking random tokens are sampled and
d is the hidden size of the low-level module. The replaced with a special [MASK] token. (2) To-
output of the discriminator can be represented as: ken Deletion randomly deletes several tokens in
the document. (3) Token Replacement randomly
H = bi-LSTM(S) replace some tokens with other tokens in the vo-
h = Max-Pooling(H) cabulary. (4) Sentence Permutation randomly swap
(4) some tokens in the sentence.
ĥ = MLP(h) Next Sentence Generation: One of the prob-
ŷ = Softmax(ĥ) lems of the DAE is that the input is always the cor-
rupted sentence, which is not the case during fine-
h ∈ Rd is a pooled representations of the discrim- tune, the pretrain-finetune discrepancy may hurt the
inator for classification. ŷ is the language distri- performance of the downstream tasks (Yang et al.,
bution in RC where C is the number of languages. 2019). Similar to Kiros et al. (2015) and Dong et al.
For the discriminator, the target is to maximize the (2019), we sample a consecutive segment in the
probability of the corresponding language while text and divide it into two parts, we treat the first
the low-level module (generator) tries to minimize parts as the source and the second part as the target.
it. Therefore, they form an adversarial training The objective is to generate the second part based
objective that the low-level module must produce on the first part.
2265
3.5 Question Generation Fine-tuning QG Pre-train
After pre-training, we suppose the low-level mod- Train Dev/Test Name(Size)
ule of our model has learned the multi-lingual lin- En 86,635 8,965/8,964 enwiki(13.6Gb)
guistic information. Then the fine-tuning objective Zh 180,000 20,000/24,962 zhwiki(1.3Gb)
is to adjust the high-level module for question gen- Ko 60,407 5,774/3,898 kowiki(608Mb)
eration. Therefore, in this phase, we fix the low- Fr 20,731 3,188/2,189 frwiki(4.0Gb)
level module, i.e. the word embedding, LSTM, and
Hi 4,000 1,300/1,255 hiwiki(395Mb)
output projection linear layer, and only update the
parameter of the high-level module. Table 1: The statistics of the multi-lingual pre-training
data and question generation data.
4 Experiments
4.1 Dataset more than 5 questions for each paragraph. Since
The question generation datasets are sometimes we did not give the specific answer candidates for
directly derived from the corresponding question each paragraph, the annotators were encouraged
answering datasets. In the current question an- to ask more general and comprehensive questions.
swering application, most multi-lingual datasets We also ask other volunteers to check the quality
are automatically derived by translating from En- and remove the questions that are either unanswer-
glish SQuAD (Asai et al., 2018). However, it may able or contain grammar errors. Finally, we obtain
reduce the multi-lingual QG tasks to translation 224,962 question-paragraph pairs. We randomly se-
tasks if we use these datasets. Therefore, we con- lect 180k of them as the training data, 20k samples
sider four different language QG datasets that are for development, and the rest 24,962 for testing.
developed by the specific language speakers. We name it LAB (Learning to Ask on Baike).
We adopt the 2020-05-20 data dumps of the
• English (En) We use the SQuAD (Rajpurkar Wikipedia4 in the corresponding language as the
et al., 2016) as the English question generation pre-training data. The details of the training data
dataset. It is a standard machine reading com- are shown in Table 1.
prehension data consists of nearly 100k human-
labeled questions from Wikipedia. 4.2 Implementation Details
• Korean (Ko) We use the Korquad1.0 (Lim et al.,
In all experiments, we tokenize the text with sen-
2019) as the Korean QG data. It consists of more
tencepiece (Kudo and Richardson, 2018). For all
than 70,000 human-generated question-answer
languages datasets, we set the vocabulary size to
pairs on Korean Wikipedia articles.
30,000. We use the Adam (Kingma and Ba, 2014)
• French (Fr) We adopt the French SQuAD-style
optimizer with 5k warm-up steps and linearly de-
dataset (d’Hoffschmidt et al., 2020) consisting of
cay the learning rate. β1 , β2 , was set to 0.9, 0.99
more than 25k human-curated French questions.
and 10−6 , respectively. For both pre-training and
• Hindi (Hi) HiQuAD (Kumar et al., 2019) is
fine-tuning, the max learning rate was set to 10−4 .
a specific Hindi QG dataset containing 6,555
The batch size was 256 during pre-training and
question-answer pairs. It was derived from the
64 during fine-tuning. We limit the max sequence
Hindi storybook.
length to 512. For the adversarial decoupling mod-
Since the size of the QG dataset except English is ule training, following previous works of genera-
comparative small, so we propose a new large-scale tive adversarial networks (Goodfellow et al., 2014;
QG dataset created by humans on Chinese (Zh). Salimans et al., 2016), the update rate for discrimi-
First of all, we collect nearly 3.5m passages from nator and generator was set to 1:10. For each of the
Baike3 , a Chinese Wikipedia-like encyclopedia. To 4 noising strategies in pre-training, we set the sam-
increase the diversity of the selected paragraphs, ple probability to 0.1. Similar with Scialom et al.
we cluster the passages based on the bag-of-words, (2019) we do not provide the answer and direcetly
then we use Ward (Ward Jr, 1963) algorithm to generate questions based on the context. We use
select the centroid in each cluster, which result in three types of models:
nearly 100k passages. We ask volunteers to ask no LALMshare is the shared language-agnostic lan-
3 4
[Link] [Link]
2266
Transformer NQG++ Multi-BERT CLQG XNLG LALMshare LALMbase LALMlarge LALMlarge +ADM
BLEU-4 14.03 15.09 17.19 17.63 19.98 20.96 21.95 23.50 24.94
En METEOR 17.62 18.04 18.38 18.91 20.24 20.23 21.30 22.15 23.28
ROUGE 40.79 40.24 44.82 43.34 46.51 47.47 48.23 50.34 51.42
BLEU-4 22.75 20.32 35.08 34.96 37.40 36.11 38.32 43.19 44.10
Zh METEOR 17.24 18.95 26.10 26.54 27.13 27.28 27.99 32.38 33.04
ROUGE 30.14 29.87 38.46 40.11 42.15 43.25 44.49 45.16 46.40
BLEU-4 7.11 7.95 10.35 8.97 - 11.93 12.19 12.58 12.93
Ko METEOR 14.30 14.81 18.10 17.22 - 19.85 20.11 20.96 21.10
ROUGE 22.17 24.13 31.28 29.34 - 34.10 34.88 34.79 35.02
BLEU-4 4.48 5.03 8.95 10.18 12.93 13.38 13.95 14.87 15.28
Fr METEOR 13.05 13.19 15.91 16.28 18.37 17.75 18.20 18.84 19.92
ROUGE 32.17 31.66 39.34 41.23 40.96 41.15 42.80 43.11 44.51
BLEU-4 9.77 10.10 23.15 20.24 - 30.35 32.21 34.02 35.19
Hi METEOR 23.85 24.32 30.29 29.15 - 33.80 34.22 35.97 36.25
ROUGE 33.16 34.91 41.06 40.64 - 48.82 49.14 50.94 51.23
Table 2: Main result of the multi-language QG. LALMshare is similar with previous multi-lingual model that the
parameters are shared across all languages. ADM represents the model trained with adversarial decoupling module.
guage model. It is similar with the proposed model in sequence-to-sequence learning. For each lan-
but has no specific low-level LSTM for each lan- guage, we train the correspondent Transformer
guage. That is, the low-level and high-level pa- model based on its training data. We set dropout
rameters are both shared across different languages. ratio to 0.4 to prevent overfitting.
The hidden size was set to 768 and the layer size NQG++ (Zhou et al., 2017) is a popular neural
was set to 12, and each layer consists of 12 heads. QG model based on LSTM. It is enhanced with
We set the shared vocabulary size to 100,000. attention and copy mechanism5 .
LALMbase is the base version of our model. Multi-BERT (Devlin et al., 2018) is a multi-
It has the same hidden size as LALMshare . The lingual extension to the original BERT model. It
low-level module was single layer uni-directional was trained on the multi-lingual wikipedia. All
LSTM with hidden size 768. LALMbase has nearly the language shares the same vocabulary. We
138m parameters, where nearly half of them are adopt the way same with Rönnqvist et al. (2019)
low-level language understanding parameters. to extend BERT to language generation task.
LALMlarge is the large version of our proposed CLQG (Kumar et al., 2019) is a cross-lingual
model. The hidden size, layer size, and head size QG method based on Transformer. It is pre-
were set to 1024,24,16, respectively. The low- trained by denoising autoencoders along with
level module was two-layer uni-directional LSTMs. back-translation. We use the public implemen-
LALMlarge has 548m parameters, where nearly a tation6 and adopt the same word tokenization as
quarter of them are low-level module’s parameters. well as pre-training data as our model.
XNLG (Chi et al., 2019) is a multi-lingual lan-
4.3 Criterion: guage generation model that transfers monolin-
Following previous works of QG (Zhou et al., 2017; gual supervision to all pre-trained languages. It
Chen et al., 2019), we adopt three widely used auto- was trained with English, Chinese and French
matic metrics for evaluation: BLEU, Meteor and datasets. We use their public pre-trained models7
Rouge-L, which measure the n-gram similarities and fine-tune on the three QG dataset.
between the generated questions and real questions.
4.5 Multi-Lingual Question Generation
4.4 Baselines To evaluate the multi-lingual question generation
We adopt 5 baseline methods for comparison. ability of the proposed methods, we assemble all
5
[Link]
Transformer (Vaswani et al., 2017; Scialom 6
[Link]
7
et al., 2019) is the most widely used architecture [Link]
2267
BLEU-4 ROUGE 4.6 Human Evaluation
Zh Zh
The automatic metrics are sometimes biased to-
F F
3 7 3 7 ward a specific attribute of the generated question
P P
(Hosking and Riedel, 2019). So we conduct hu-
3 38.32 36.03 3 44.49 41.73
man qualitative evaluation of the generated outputs.
7 – 34.22 7 – 40.12
We consider three aspects of the generated ques-
En En
tions: Fluency: Whether the generated questions
F F
3 7 3 7 are well-posed and natural, in terms of both gram-
P P
mar and semantic. Answerable: Whether the gen-
3 21.95 20.61 3 48.23 47.15
erated questions could be answered by the context
7 - 17.93 7 - 45.02
paragraph. Significance: Whether the generated
Table 3: Multi-lingual and mono-lingual results for question is just a simple syntactical transformation
LALMbase . P denotes the pre-training and F denotes of the paragraph sentence or trivial one that seems
the fine-tuning, where 3denote the multi-lingual while unlikely asked by human.
7denotes the mono-lingual training. For example, the We randomly sample 50 generated questions
upper right cell in each table denotes pre-training with from English and 50 from Chinese and ask three
multi-lingual but finetuning with mono-lingual.
volunteers to evaluate the sample quality. The re-
sult is shown in Table 5. The result shows our
proposed model is also excels at human evalua-
tion, especially for significance, which is some-
QG data and train the LALM thereof. For Trans-
times regarded as the most important factor in QG
former and NQG++, we initialize the word embed-
(Graesser et al., 2010). We also showcase some
dings by fasttext multilingual word embeddings
outputs of our model in Table 4. We can see that
(Grave et al., 2018). The result is shown in Table 2.
LALM could generate fluent and sound questions.
We can see from the table that our model ex-
cels at multi-lingual QG, achieving significant im- 4.7 Multi-Lingual v.s. Mono-Lingual
provement over previous methods in all languages. Kumar et al. (2019) have found that in QG the per-
Compared with other architectures such as Trans- formance of Hindi could be improved by training
former, we explicitly separate the low-level and the with additional English data. In this section, we
high-level module in the proposed model and use evaluate whether the multi-lingual is superior to
adversarial networks to decouple them. Therefore, the mono-lingual QG. We focus on two aspects:
the shared high-level module is encouraged to learn (1) Pre-training. In contrast to the proposed multi-
more common representations across different lan- lingual pre-training, we adopt the mono-lingual
guages, which is more transferable and benefits the pre-training where we only pretrain on specific lan-
downstream QG task a lot. guages8 and fine-tune the QG models in the same
language. (2) Fine-tuning. Different from the
Besides, we can see that if we don’t explic- setup in Sec. 4.5 where we aggregate all languages
itly separate the low and high-level parameters QG data for training, we only fine-tune the model
(LALMshare ), the result drops a lot. We hypothesis on specific language.
that different languages have different low-level We experiment on English and Chinese with the
language information, such as lexical, syntactical, LALM base model. The BLEU-4 and ROUGE-L
etc. Embedding all language processing procedures scores are shown in Table 3. It is clear that for
into a single model may make the model hard to both pre-training and fine-tuning, the multi-lingual
discriminate the language-specific information. training improves the model a lot. Moreover, the
Besides, the model trained with the adversarial multi-lingual plays a more important role in pre-
decoupling module achieves further improvement, training than in fine-tuning. We suppose that dur-
the ADM may impose an implicit regularization ing pre-training, multiple languages perform a type
on the low-level module to make the representa- of regularization on the shared high-level module,
tions more abstract, and therefore encourage the while in fine-tuning the language-dependent super-
high-level module to learn more common represen- 8
Therefore, we omit the adversarial decoupling module
tations (Chen et al., 2017; Liu et al., 2017). since it only takes effect on multi-lingual learning.
2268
Table 4: Some generated cases of the proposed model.
English
Context: The United Methodist Church opposes conscription as incompatible with the teaching of Scripture. Therefore, the
Church supports and extends its ministry to those persons who conscientiously oppose all war, or any particular war, and
who therefore refuse to serve in the armed forces or to cooperate with systems of military conscription. However, the United
Methodist Church also supports and extends its ministry to those persons who conscientiously choose to serve in the armed
forces or to accept alternative service. The church also states that ”as Christians they are aware that neither the way of
military action, nor the way of inaction is always righteous before God.”
Original: The Church supports those persons who conscientiously oppose what?
LALM: what does the church states after they oppose the construction ?
Chinese
Context: 电桥平衡#四个电阻R0、R1、R2、Rx连成四边形,称为电桥的四个臂。四边形的一个对角线连有检流
计,称为“桥”;四边形的另一对角线接上电源,称为电桥的“电源对角线”。E为线路中供电电源,学生实验用双路直
流稳压电源,电压可在0-30V之间调节。R保护为较大的可变电阻,在电桥不平衡时取最大电阻作限流...
Original: 什么是电桥平衡?
LALM: 电桥平衡有什么用?
French
Context: Le seul quartier d’habitation à avoir été fouillé est situé sur le site du Merkes, à l’est de la Voie processionnelle et
du complexe sacré, entre les anciens quartiers de Ka-dingirra, Eridu et Shuanna. Sa voirie est caractérisée par des rues
étroites approximativement rectilignes et se coupant quasiment à angles droits. Il s’agit peut-être de l’héritage d’un ancien
plan orthogonal planifié qui a été altéré à la suite de remaniements de constructions, courants en raison de l’altération
rapide des constructions en briques crues qui doivent régulièrement être restaurées.
Original: En quoi sont fabriquées les habitations ?
LALM: Quelles sont les caractéristiques de la route ?
Korean
Context: ᄎ ᆼᄍ
ᅵ ᅡ
ᆼ(藏)ᄀ ᅩᄋ ᆫᄋ
ᅯ ᅵᄅ ᅡᄀ ᅩᄃ ᅩᄇ ᆯᄅ
ᅮ ᅵᄂ ᆫᄐ
ᅳ ᅵᄇ ᅦ트ᄀ ᅩᄋᆫᄋ
ᅯ ᆫᄃ
ᅳ ᆼᄋ
ᅩ ᅡᄉ ᅵᄋ ᅡᄋ ᅦᄋ ᅱᄎ ᅵᄒᆫᄂ
ᅡ ᆲᄀ
ᅥ ᅩᄂ ᇁᄋ
ᅩ ᆫᄀ
ᅳ ᅩᄋ ᆫᄋ
ᅯ ᅵᄃ ᅡ. 티베ᄐ ᅳ자ᄎ ᅵ구ᄋ ᆨ
ᅧ
ᅪᄌ
ᄀ ᆼᄀ
ᅮ ᆨᄎ
ᅮ ᆼᅡ
ᅵ ᄒ이ᄉᆼ(海省), ᄀ
ᅥ ᅳ리ᄀ ᅩᄋ ᆫᄃ
ᅵ ᅩᄏ ᅡᄉ ᅲᄆ ᅵᄅ ᅳᄋ ᅦᄀ ᆯᄎ
ᅥ ᅧᄋᆻᄂ
ᅵ ᆫᄐ
ᅳ ᅵᄇ ᅦᄐ ᅳᄀ ᅩᄋ ᆫᄋ
ᅯ ᆫᄂ
ᅳ ᆷᄇ
ᅡ ᆨ 1000km, ᄃ
ᅮ ᆼᄉ
ᅩ ᅥ 2500kmᄋ ᅦᄈ ᆮᄋ
ᅥ ᅥᄋ ᆻ
ᅵ
ᅳᄆ
ᄋ ᅧ, 그ᄑ ᆼᄀ
ᅧ ᆫᅩ
ᅲ ᄂᄋ
ᇁ ᅵᄂ ᆫ 4500 ᄆ
ᅳ ᅵᄐ ᅥᄀ ᅡᄂ ᆷᄂ
ᅥ ᆫᄃ
ᅳ ᅡ. ’ᄉ ᅦᄀ ᅨᄋ ᅴᄌ ᅵᄇᆼ’ᄋ
ᅮ ᅳᄅ ᅩᄇ ᆯᄅ
ᅮ ᆯᄆ
ᅵ ᅡ
ᆫᄏ ᆷᄉ
ᅳ ᅦ계에ᄉ ᅥᄀ ᅡᄌ ᅡ
ᆼᄂ ᇁᄀ
ᅩ ᅩᄏ ᅳᄆ ᅧᄆ ᆫᄌ
ᅧ ᆨᄋ
ᅥ ᆫᄋ
ᅳ ᆨ 250ᄆ
ᅣ ᅡ
ᆫ
ᆼᄇ
ᄑ
ᅧ ᅡ
ᆼᄏ ᆯᄅ
ᅵ ᅩ미ᅥ ᄐ나ᄃ ᅬ
ᆫ다. 이ᄀ ᅩᄋ ᆫᄋ
ᅯ ᆫᄋ
ᅳ ᆫᄃ
ᅵ ᅩ-ᄒ ᅩᄌ ᅮᄑ ᆯᄅ
ᅳ ᅦ이ᄐ ᅳᅪᄋᄋ ᅲᄅ ᅡᄉ ᅵ아ᄑ ᆯᄅ
ᅳ ᅦ이트가ᄉ ᆫᄉ
ᅵ ᆼᄃ
ᅢ ᅢ에ᄎ ᆼᄃ
ᅮ ᆯᄒ
ᅩ ᅡᄆ ᅧᄉ ᆼᄉ
ᅢ ᆼᄃ
ᅥ ᅬᄋ ᆻᄋ
ᅥ ᅳᅧ ᄆ
ᅳᄀ
ᄀ ᅪᄌ ᆼᄋ
ᅥ ᆫᅵ
ᅳ ᄌᄀᆷᄃ
ᅳ ᅩᄌ ᆫᄒ
ᅵ ᆼᄃ
ᅢ ᅬᄀ ᅩᄋ ᆻᄃ
ᅵ ᅡ. ᄋ ᅵᄀ ᅩᄋ ᆫᄋ
ᅯ ᆫᄉ
ᅳ ᆫᄆ
ᅡ ᆨᄀ
ᅢ ᅪᄉ ᅩᄀᆷᄒ
ᅳ ᅩᄉ ᅮᄀ ᅡᄇ ᆫᄑ
ᅮ ᅩᄒ ᆫᄀ
ᅡ ᆫᄋ
ᅩᄋ
ᅯ ᅴᄀ ᆫᄌ
ᅥ ᅩᄉ ᅳᄐ ᆸᄌ
ᅦ ᅵᄃ ᅢᄅᆯᄒ
ᅳ ᆼᅥ
ᅧ ᆼ
ᄉ하ᄀ ᅩᄋ ᆻᄃ
ᅵ ᅡ.
ᆫᄒ
ᄒ
ᅡ ᅢᄑ ᆼᄀ
ᅧ ᆫᄀ
ᅲ ᅡ
ᆼᄉ ᅮ랴
ᆼᄋ ᆫ 100mmᄋ
ᅳ ᅦᄉ ᅥ 300mmᄅ ᅩ, 가
ᆼᄉ ᅮ랴
ᆼᄋ ᅴ대ᄇ ᅮᄇ ᆫᄋ
ᅮ ᆫᄋ
ᅳ ᅮᄇ ᆨᄋ
ᅡ ᆯᄋ
ᅳ ᅵᄅ ᆫᄃ
ᅮ ᅡ. 유ᄆ ᆨᄆ
ᅩ ᆫᄃ
ᅵ ᆯᄋ
ᅳ ᆫᄀ
ᅳ ᅩᄋ ᆫᄋ
ᅯ ᅴᄂᆷᄇ
ᅡ ᅮᄆ ᆾᄃ
ᅵ ᆼᄇ
ᅩ ᅮ
ᆼᄀ
ᄀ
ᅧ ᅨᄋ ᅴᅡ ᆫ
해 ᄒ 6개ᄋᆯᄀ
ᅯ ᅡ랴
ᆼᄉ ᅥᄅ ᅵᄀ ᅡᄂ ᅢᄅ ᅵᄂ ᆫᄆ
ᅳ ᆨᄎ
ᅩ ᅩ지ᄋ ᅦᄉ ᅥᄋ ᅲᄆᆨᄉ
ᅩ ᆼᄒ
ᅢ ᆯᄋ
ᅪ ᆯᄋ
ᅳ ᅲᄌ ᅵᄒ ᅡᄀ ᅩᄋ ᆻᄃ
ᅵ ᅡ.
Original: ᄐ ᅵᄇ ᅦ트고ᄋ ᆫᄋ
ᅯ ᅴᄆ ᆫᄌ
ᅧ ᆨᄋ
ᅥ ᆫ?
ᅳ
LALM: ᄐ ᅵ베트ᄀ ᅩᄋᆫᄋ
ᅯ ᆫᄋ
ᅳ ᅥ디ᄋ ᅦᄋ ᆻᄉ
ᅵ ᆸᄂ
ᅳ ᅵᄁ ᅡ?
2269
En Zh Hi
B4 M R B4 M R B4 M R
LALM 14.35 17.41 33.52 22.25 21.43 37.94 12.92 24.19 33.10
LALM+ADM 15.94 18.24 36.24 24.10 22.05 38.78 14.13 23.77 34.24
LALM+ADM+Pre-train 21.95 21.30 48.23 38.32 27.99 44.49 30.35 34.80 48.82
Table 6: Ablation study of the pre-training. The three models are fine-tuned on multi-lingual data.
B1 B2 B3 B4 M R Acknowledgments
F 43.07 31.04 23.58 17.74 18.06 22.44 We thank the anonymous reviewers for their in-
Zh
Z 26.55 18.26 12.10 10.94 11.94 15.89 sightful comments. And we appereate the dedi-
F 25.20 15.34 10.71 5.07 14.35 16.42 cated labeling efforts contributed by the annoators,
Kr
Z 20.55 11.17 8.32 5.95 13.32 16.77 which makes the large-scale Chinese QG datasets
F 31.31 14.91 10.46 5.54 9.63 23.02 avaliable for the community.
Fr
Z 25.58 13.33 12.49 6.32 11.06 15.22
F 30.15 20.42 12.30 9.03 23.47 32.84 References
Hi
Z 24.10 15.77 12.54 10.89 26.42 33.01
Akari Asai, Akiko Eriguchi, Kazuma Hashimoto, and
Yoshimasa Tsuruoka. 2018. Multilingual extractive
Table 7: Zero-shot multi-lingual evaluation. F denotes reading comprehension by runtime machine transla-
the performance of NQG++ model, and Z denotes zero- tion. arXiv preprint arXiv:1809.03275.
shot result where we fine-tune LALM base model on
SQuAD and directly evaluate on other datasets. Xinchi Chen, Zhan Shi, Xipeng Qiu, and Xuan-Jing
Huang. 2017. Adversarial multi-criteria learning for
chinese word segmentation. In Proceedings of the
are only 4,000 training instances. Nevertheless, 55th Annual Meeting of the Association for Compu-
tational Linguistics (Volume 1: Long Papers), pages
when trained with the adversarial decoupling mod-
1193–1203.
ule, our model could achieve consistent improve-
ment, demonstrating that the ADM is good at multi- Yu Chen, Lingfei Wu, and Mohammed J Zaki. 2019.
lingual transfer learning. Reinforcement learning based graph-to-sequence
model for natural question generation. arXiv
preprint arXiv:1908.04942.
5 Conclusion
Zewen Chi, Li Dong, Furu Wei, Wenhui Wang, Xian-
In this paper, we propose a language-agnostic lan- Ling Mao, and Heyan Huang. 2019. Cross-lingual
guage model to deal with the multi-lingual question natural language generation via pre-training. arXiv
generation. The model consists of the low-level preprint arXiv:1909.10481.
and the high-level module to explicitly represent
Alexis Conneau, Ruty Rinott, Guillaume Lample, Ad-
the language-dependent and language-independent ina Williams, Samuel Bowman, Holger Schwenk,
information, respectively. We operate the attention and Veselin Stoyanov. 2018. Xnli: Evaluating cross-
mask matrix to fit our model to the sequence to lingual sentence representations. In Proceedings of
sequence learning. We propose an adversarial train- the 2018 Conference on Empirical Methods in Natu-
ral Language Processing, pages 2475–2485.
ing mechanism to decouple the two-level modules,
making the low-level module contains more ab- Yiming Cui, Wanxiang Che, Ting Liu, Bing Qin, Shi-
stractive representations and the high-level module jin Wang, and Guoping Hu. 2019. Cross-lingual
language-agnostic. We also proposed a large-scale machine reading comprehension. In Proceedings of
the 2019 Conference on Empirical Methods in Nat-
Chinese QG data containing more than 220k ques- ural Language Processing and the 9th International
tions. Experiments on five languages demonstrate Joint Conference on Natural Language Processing
our model achieves significant improvements over (EMNLP-IJCNLP), pages 1586–1595.
previous methods in multi-lingual QG. For future
Beth Davey and Susan McBride. 1986. Effects
work, we would like to apply our proposed model of question-generation training on reading com-
to other multi-lingual tasks such as summarization prehension. Journal of Educational Psychology,
and question answering. 78(4):256.
2270
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Hafedh Hussein, Mohammed Elmogy, and Shawkat
Kristina Toutanova. 2018. Bert: Pre-training of deep Guirguis. 2014. Automatic english question genera-
bidirectional transformers for language understand- tion system based on template driven scheme. IJCSI,
ing. arXiv preprint arXiv:1810.04805. 11(6):45.
Martin d’Hoffschmidt, Maxime Vidal, Wacim Belb- Diederik P. Kingma and Jimmy Ba. 2014. Adam: A
lidia, and Tom Brendlé. 2020. Fquad: French method for stochastic optimization. ICLR.
question answering dataset. arXiv preprint
arXiv:2002.06071. Ryan Kiros, Yukun Zhu, Russ R Salakhutdinov,
Richard Zemel, Raquel Urtasun, Antonio Torralba,
and Sanja Fidler. 2015. Skip-thought vectors. In
Kaustubh D Dhole and Christopher D Manning.
Advances in neural information processing systems,
2020. Syn-qg: Syntactic and shallow seman-
pages 3294–3302.
tic rules for question generation. arXiv preprint
arXiv:2004.08694. Taku Kudo and John Richardson. 2018. Sentencepiece:
A simple and language independent subword tok-
Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xi- enizer and detokenizer for neural text processing.
aodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, arXiv preprint arXiv:1808.06226.
and Hsiao-Wuen Hon. 2019. Unified language
model pre-training for natural language understand- Vishwajeet Kumar, Nitish Joshi, Arijit Mukherjee,
ing and generation. In Advances in Neural Informa- Ganesh Ramakrishnan, and Preethi Jyothi. 2019.
tion Processing Systems, pages 13042–13054. Cross-lingual training for automatic question gen-
eration. In Proceedings of the 57th Annual Meet-
Xinya Du, Junru Shao, and Claire Cardie. 2017. Learn- ing of the Association for Computational Linguistics,
ing to ask: Neural question generation for reading pages 4863–4872.
comprehension. pages 1342–1352.
Guillaume Lample and Alexis Conneau. 2019. Cross-
Nan Duan, Duyu Tang, Peng Chen, and Ming Zhou. lingual language model pretraining. arXiv preprint
2017. Question generation for question answering. arXiv:1901.07291.
In EMNLP, pages 866–874.
Mike Lewis, Yinhan Liu, Naman Goyal, Mar-
jan Ghazvininejad, Abdelrahman Mohamed, Omer
Xiangyu Duan, Mingming Yin, Min Zhang, Boxing
Levy, Ves Stoyanov, and Luke Zettlemoyer. 2019.
Chen, and Weihua Luo. 2019. Zero-shot cross-
Bart: Denoising sequence-to-sequence pre-training
lingual abstractive sentence summarization through
for natural language generation, translation, and
teaching generation and attention. In Proceedings of
comprehension. arXiv preprint arXiv:1910.13461.
the 57th Annual Meeting of the Association for Com-
putational Linguistics, pages 3162–3172. Seungyoung Lim, Myungji Kim, and Jooyoul Lee.
2019. Korquad1. 0: Korean qa dataset for ma-
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, chine reading comprehension. arXiv preprint
Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron arXiv:1909.07005.
Courville, and Yoshua Bengio. 2014. Generative ad-
versarial nets. In Advances in neural information Jiahua Liu, Yankai Lin, Zhiyuan Liu, and Maosong
processing systems, pages 2672–2680. Sun. 2019. Xqa: A cross-lingual open-domain ques-
tion answering dataset. In Proceedings of the 57th
Art Graesser, Yasuhiro Ozuru, and Jeremiah Sullins. Annual Meeting of the Association for Computa-
2010. What is a good question? tional Linguistics, pages 2358–2368.
Edouard Grave, Piotr Bojanowski, Prakhar Gupta, Ar- Pengfei Liu, Xipeng Qiu, and Xuan-Jing Huang. 2017.
mand Joulin, and Tomas Mikolov. 2018. Learning Adversarial multi-task learning for text classifica-
word vectors for 157 languages. In Proceedings tion. In Proceedings of the 55th Annual Meeting of
of the International Conference on Language Re- the Association for Computational Linguistics (Vol-
sources and Evaluation (LREC 2018). ume 1: Long Papers), pages 1–10.
Tom Hosking and Sebastian Riedel. 2019. Evaluating Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and
rewards for question generation models. In NAACL, Percy Liang. 2016. Squad: 100,000+ questions for
pages 2278–2283. machine comprehension of text. In Proceedings of
2271
the 2016 Conference on Empirical Methods in Natu- Machine comprehension by text-to-text neural ques-
ral Language Processing, pages 2383–2392. tion generation. arXiv preprint arXiv:1705.02012.
Adam Roberts, Colin Raffel, and Noam Shazeer. 2020. Qingyu Zhou, Nan Yang, Furu Wei, Chuanqi Tan,
How much knowledge can you pack into the pa- Hangbo Bao, and Ming Zhou. 2017. Neural ques-
rameters of a language model? arXiv preprint tion generation from text: A preliminary study. In
arXiv:2002.08910. NLPCC.
Samuel Rönnqvist, Jenna Kanerva, Tapio Salakoski, Junnan Zhu, Qian Wang, Yining Wang, Yu Zhou, Jiajun
and Filip Ginter. 2019. Is multilingual bert flu- Zhang, Shaonan Wang, and Chengqing Zong. 2019.
ent in language generation? arXiv preprint Ncls: Neural cross-lingual summarization. arXiv
arXiv:1910.03806. preprint arXiv:1909.00156.
2272