●
●
●
●
●
●
●
●
○ Occurrence-based representation
■ textual features are computed on the frequencies of occurrences of the
main textual units (e.g., words, n-grams, phrases) in the input text
○ Distributed vector representations of text
■ high-dimensional text representation extracted by means of neural network
training
●
●
●
○ Occurrence-based representation
■ textual features are computed on the frequencies of occurrences of the
main textual units (e.g., words, n-grams, phrases) in the input text
○ Distributed vector representations of text
■ high-dimensional text representation extracted by means of neural network
training
●
●
●
○ Number of words in the dictionary >> Number of documents
●
○ Each document is likely to include a limited number of words in the dictionary
●
○ Counteract data sparsity
●
●
●
●
●
●
○ Each word is representative of a local pattern
○ Term frequency positively contributes to the relative word importance
●
○ Word occurrences are spread over all documents
○ Document frequent negatively contributes to the relative word importance
●
●
●
○ Heavy tailed or skewed distributions of tf-idf values may come out
●
○ Each of them contains slightly different components and parameters
●
○ Avgdl: average document length (expressed as number of words) over all the
documents in the collection
○ k1: term frequency saturation (free parameter)
○ b: penalty score associated with the document length (free parameter)
●
○ To scale the skewed weight distributions
●
○ To bound the range of variation of the frequency values
●
○ To set a unique free parameter K
●
○ [Link]
[Link]
●
○ [Link]
●
●
●
○
●
○
●
○ Mainly unsupervised
●
●
Rob Churchill and Lisa Singh. 2022. The Evolution of Topic Modeling. ACM Comput. Surv. 54, 10s, Article 215 (January 2022), 35 pages. [Link]
●
●
○ Vector representations of text allow clustering text into homogenous groups
○ Occurrence-based representations are high-dimensional can be reduced to low-
dimensional spaces
●
●
●
○ The underlying concepts and the corresponding weights are derived from the
Singular Value Decomposition
●
○ Derive document-topics and topic-terms probabilities distributions using a
generative process based on statistical Bayesian inference
●
●
○ Matrix factorization
●
(right) eigenvector eigenvalue
●
●
Christopher D. Manning and Pandu Nayak. Introduction to Information Retrieval. CS276
only has a non-zero solution if
Christopher D. Manning and Pandu Nayak. Introduction to Information Retrieval. CS276
●
A =USV T
MM MN V is NN
Christopher D. Manning and Pandu Nayak. Introduction to Information Retrieval. CS276
A =USV T
MM MN V is NN
●
●
●
Christopher D. Manning and Pandu Nayak. Introduction to Information Retrieval. CS276
A =USV T
MM MN V is NN
s i = li
S = diag(s 1...s r )
σi+1 <= σi
Christopher D. Manning and Pandu Nayak. Introduction to Information Retrieval. CS276
●
○ A lower dimensional representation of the occurrence-based text representation
Ak = min
X:rank( X)=k
A- X F
●
●
●
min A- X F = A- Ak F
= s k+1
X:rank( X)=k
●
○ All the other are reduced to zero
●
MN Mk kk
kN
Christopher D. Manning and Pandu Nayak. Introduction to Information Retrieval. CS276
∑
A =
U
VT
∑’
A’ =
U
VT
∑
A =
U
VT
K selected topics Topic relevance
∑’
Documents
A’ =
U
VT
Words K selected topics
●
●
○ Maps terms and documents to a lower dimensional representation
○ Preserve pairwise word associations conveyed by the occurrence-based model
min A- X F = A- Ak F
= s k+1
X:rank( X)=k
●
●
∑ ∑
●
●
○ The i-th concept corresponding to σi is likely to be more relevant
than those associated with σi+1
●
○ We multiply the diagonal matrix ∑’ by a specific orthogonal eigenvector
in VT
●
○ We multiply the specific orthogonal eigenvector in U by the diagonal
matrix ∑’
Christopher D. Manning and Pandu Nayak. Introduction to Information Retrieval. CS276
●
●
○ define σk as the smallest singular value that is above the half of
the highest one (i.e., σk > σ1/2)
●
●
○ the matrix is sparse
○ we consider only the k << M singular vectors (reduced SVD)
■ Typical situation while coping with textual data
●
○ MATLAB ([Link]
○ SK-Learn ([Link]
○ Hadoop Spark ([Link]
○ R ([Link]
r1 c1 c2
r2
●
○ Heuristic approach -> K = 2
●
○ Topic #1
■ r1 x c1 = -4.31
○ Topic #2
■ r2 x c1 = 6.85
r1
c2
●
○ d1 = [-4.31, 6.85]
○ d2 = [-4.31, 6.85]
○ d3 = [-7.14, -2.72]
○ d4 = [-7.14, -2.72]
○ d5 = [-6.87, -3.16]
●
●
○ W1 : 5.984 (= r1 x c2)
○ W2 : 5.11
○ W3 : 5.11
○ W4 : -3.16
○ W5 : -3.92
○ W6: -1.96
●
●
●
●
○ They range over multiple topics in different proportions
Generating Summary Keywords for Emails Using Topics. Mark Dredze, Hanna M. Wallach, Danny Puller, , Fernando Pereira. IUI’08. ACM. 2008
●
●
●
○ The n-th iteration does not produce any significant improvement
compared to the (n-1)-th iteration
David M. Blei, Andrew Y. Ng, Micheal I. Jordan. Latent Dirichlet Allocation. Journal of Machine Learning Research 3 (2003) 993-1022
●
●
○ by sampling a topic from a document-specific distribution over
topics and
○ by sampling a word from the distribution over words that
characterizes that topic
●
Generating Summary Keywords for Emails Using Topics. Mark Dredze, Hanna M. Wallach, Danny Puller, , Fernando Pereira. IUI’08. ACM. 2008
●
●
●
●
Generating Summary Keywords for Emails Using Topics. Mark Dredze, Hanna M. Wallach, Danny Puller, , Fernando Pereira. IUI’08. ACM. 2008
●
○ E.g., the Gibbs-Expectation Maximization (EM) algorithm is commonly used to do
statistical inference of the posterior distribution of the latent variable for a given
corpus
○ Gibbs-EM alternates between optimizing α and β and sampling a topic
assignment for each word in the corpus from the distribution over topics for that
word, conditioned on all other variables
D. M. Blei, A. Y. Ng, and M. I. Jordan. Latent dirichlet allocation. Journal of machine Learning research, 3(Jan):993–1022, 2003.
●
○ A topic is chosen accordingly to the document-topic distribution
●
David M. Blei, Andrew Y. Ng, Micheal I. Jordan. Latent Dirichlet Allocation. Journal of Machine Learning Research 3 (2003) 993-1022
●
○ the corpus
○ the documents
○ the terms
●
○ the same document can be described by multiple topics
●
●
David M. Blei, Andrew Y. Ng, Micheal I. Jordan. Latent Dirichlet Allocation. Journal of Machine Learning Research 3 (2003) 993-1022
•
•
• α corresponds to topics-per-document ratio
• Setting α higher results in more topics per document
• β corresponds to words-per-topic ratio
• Setting β lower results in fewer words per topic
D. M. Blei, A. Y. Ng, and M. I. Jordan. Latent dirichlet allocation. Journal of machine Learning research, 3(Jan):993–1022, 2003.
D. M. Blei, A. Y. Ng, and M. I. Jordan. Latent dirichlet allocation. Journal of machine Learning research, 3(Jan):993–1022, 2003.
●
Generating Summary Keywords for Emails Using Topics. Mark Dredze, Hanna M. Wallach, Danny Puller, , Fernando Pereira. IUI’08. ACM. 2008
●
●
○ integrating over the whole distribution θ
○ summing over all topics z
○ taking the product of the marginal probabilities of each document
David M. Blei, Andrew Y. Ng, Micheal I. Jordan. Latent Dirichlet Allocation. Journal of Machine Learning Research 3 (2003) 993-1022
●
●
○ MATLAB ([Link]
○ SK-Learn
([Link]
○ Hadoop Spark
([Link]
●
●
●
●
●
○ Documents
○ Terms
○ Authors
Michal Rosen-Zvi, Thomas Griffiths, Mark Steyvers, and Padhraic Smyth. 2004. The author-topic model for authors and documents. In
Proceedings of the 20th conference on Uncertainty in artificial intelligence (UAI '04). AUAI Press, Arlington, Virginia, USA, 487–494.
●
○ Who is the most authoritative author on a given topic?
○ What are the topics covered by a given author?
○ What is the most authoritative paper of an author?
Michal Rosen-Zvi, Thomas Griffiths, Mark Steyvers, and Padhraic Smyth. 2004. The author-topic model for authors and documents. In
Proceedings of the 20th conference on Uncertainty in artificial intelligence (UAI '04). AUAI Press, Arlington, Virginia, USA, 487–494.
●
●
●
Michal Rosen-Zvi, Thomas Griffiths, Mark Steyvers, and Padhraic Smyth. 2004. The author-topic model for authors and documents. In
Proceedings of the 20th conference on Uncertainty in artificial intelligence (UAI '04). AUAI Press, Arlington, Virginia, USA, 487–494.
●
Michal Rosen-Zvi, Thomas Griffiths, Mark Steyvers, and Padhraic Smyth. 2004. The author-topic model for authors and documents. In
Proceedings of the 20th conference on Uncertainty in artificial intelligence (UAI '04). AUAI Press, Arlington, Virginia, USA, 487–494.
Modeling Documents. Amruta Joshi. Stanford University. Department of Computer Science.
Michal Rosen-Zvi, Thomas Griffiths, Mark Steyvers, and Padhraic Smyth. 2004. The author-topic model for authors and documents. In
Proceedings of the 20th conference on Uncertainty in artificial intelligence (UAI '04). AUAI Press, Arlington, Virginia, USA, 487–494.
●
●
○ Θ: probability of a given word given a topic
○ φ: probability of a topic given an author
Michal Rosen-Zvi, Thomas Griffiths, Mark Steyvers, and Padhraic Smyth. 2004. The author-topic model for authors and documents. In
Proceedings of the 20th conference on Uncertainty in artificial intelligence (UAI '04). AUAI Press, Arlington, Virginia, USA, 487–494.
●
●
●
○
●
○
■
■