0% found this document useful (0 votes)
4 views9 pages

Sim Stocks

The document presents SimStock, a novel framework that utilizes self-supervised learning and temporal domain generalization to represent stock similarities effectively. It addresses challenges such as temporal distribution shifts and ambiguity in stock classifications, demonstrating superior performance in identifying similar stocks across various benchmarks. SimStock's approach simplifies investment opportunity screening by leveraging diverse data types and adapting to the dynamic nature of stock data.

Uploaded by

kenlee.reb
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views9 pages

Sim Stocks

The document presents SimStock, a novel framework that utilizes self-supervised learning and temporal domain generalization to represent stock similarities effectively. It addresses challenges such as temporal distribution shifts and ambiguity in stock classifications, demonstrating superior performance in identifying similar stocks across various benchmarks. SimStock's approach simplifies investment opportunity screening by leveraging diverse data types and adapting to the dynamic nature of stock data.

Uploaded by

kenlee.reb
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

PDF Download

[Link]
30 December 2025
Total Citations: 3
Total Downloads: 704
.
.
Latest updates: hps://[Link]/doi/10.1145/3604237.3626888

.
.
Published: 27 November 2023
.
.
.
RESEARCH-ARTICLE

.
Citation in BibTeX format
SimStock : Representation Model for Stock Similarities

.
.
ICAIF '23: 4th ACM International
Conference on AI in Finance
YOONTAE HWANG, Ulsan National Institute of Science and Technology, Ulsan, South Korea November 27 - 29, 2023
.
NY, Brooklyn, USA
JUNHYEONG LEE, Ulsan National Institute of Science and Technology, Ulsan, South Korea

.
.
.
DAHAM KIM, Cornell University, Ithaca, NY, United States
.
SEUNGHWAN NOH, Ulsan National Institute of Science and Technology, Ulsan, South Korea
.
JOOHWAN HONG, Ulsan National Institute of Science and Technology, Ulsan, South Korea
.
YONGJAE LEE, Ulsan National Institute of Science and Technology, Ulsan, South Korea
.
.
.
Open Access Support provided by:
.
Ulsan National Institute of Science and Technology
.
Cornell University
.
ICAIF '23: Proceedings of the Fourth ACM International Conference on AI in Finance (November 2023)
hps://[Link]/10.1145/3604237.3626888
ISBN: 9798400702402
.
SimStock : Representation Model for Stock Similarities
Yoontae Hwang Junhyeong Lee Daham Kim
Ulsan National Institute of Science Ulsan National Institute of Science Cornell University
and Technology and Technology New York, Republic of Korea
Ulsan, Republic of Korea Ulsan, Republic of Korea dk753@[Link]
yoontae@[Link] [Link]@[Link]

Seunghwan Noh Joohwan Hong Yongjae Lee


Ulsan National Institute of Science Ulsan National Institute of Science Ulsan National Institute of Science
and Technology and Technology and Technology
Ulsan, Republic of Korea Ulsan, Republic of Korea Ulsan, Republic of Korea
nohseunghwan@[Link] joohwanhong@[Link] yongjaelee@[Link]

ABSTRACT Navigating a wealth of information from the stock market, includ-


In this study, we introduce SimStock, a novel framework leverag- ing fundamental, technical, and sentiment data, can often be daunt-
ing self-supervised learning and temporal domain generalization ing and labor-intensive. Fortunately, when large-scale data and
techniques to represent similarities of stock data. Our model is various data types are considered, self-supervised learning (SSL)
designed to address two critical challenges: 1) temporal distribution holds promise for discovering unique representations (i.e., low-
shift (caused by the non-stationarity of financial markets), and 2) dimensional embeddings) of stock data. In particular, due to their
ambiguity in conventional regional and sector classifications (due to versatility, these representations can be utilized in a wide range of
rapid globalization and digitalization). SimStock exhibits outstand- sub-tasks such as forecasting, classification, and anomaly detection.
ing performance in identifying similar stocks across four real-world Despite the recent successes of SSL in areas like images[7][18][38]
benchmarks, encompassing thousands of stocks. The quantitative and natural language[21][27][28][29], there has been relatively less
and qualitative evaluation of the proposed model compared to vari- focus on SSL for time series data with non-stationary characteristics,
ous baseline models indicates its potential for practical applications such as stock data.
in stock market analysis and investment decision-making. We list two major challenges when applying SSL to achieve good
representations of stock data: 1) temporal distribution shift, and
CCS CONCEPTS 2) ambiguity of regional and sector classifications. First, it has al-
ways been difficult to deal with temporal distribution shifts (or
• Computing methodologies → Artificial intelligence. non-stationarity) in stock data. For instance, the price movement
of a specific stock could vary considerably across different time
KEYWORDS periods, and it can be challenging to understand the relationships
Similar stocks, Self-supervised learning, Stock representation, Tem- between various stocks because of their dynamic interactions. Sec-
poral distribution shift, Domain generalization ondly, the traditional means of categorizing stocks based on regional
or sector classifications have become increasingly ambiguous due
ACM Reference Format: to globalization and digitalization. For instance, a company like
Yoontae Hwang, Junhyeong Lee, Daham Kim, Seunghwan Noh, Joohwan
Samsung Electronics, although listed in South Korea, operates fac-
Hong, and Yongjae Lee. 2023. SimStock : Representation Model for Stock
Similarities. In 4th ACM International Conference on AI in Finance (ICAIF
tories in various countries including the U.S., China, and Vietnam.
’23), November 27–29, 2023, Brooklyn, NY, USA. ACM, New York, NY, USA, Similarly, while Amazon is usually categorized under consumer
8 pages. [Link] goods and Tesla under automobiles, their actual business operations
span across much more diverse industry sectors. Therefore, to ef-
fectively tackle these challenges, it is essential to comprehensively
1 INTRODUCTION leverage various data sources, including stock price data, financial
Machine learning is being actively utilized in finance [22], but it is statements, and text descriptions[35].
difficult to utilize the vast and multifaceted stock data effectively. Unfortunately, most existing works in SSL have focused on
invariance[8][15]. That is, they rely on simple inductive biases
that two similar observations should yield similar outputs, and
Permission to make digital or hard copies of all or part of this work for personal or
classroom use is granted without fee provided that copies are not made or distributed
these have proven to be effective when augmenting data (mostly
for profit or commercial advantage and that copies bear this notice and the full citation for images). However, for non-stationary data, such as stocks, it
on the first page. Copyrights for components of this work owned by others than the is quite challenging to incorporate these distribution shifts into
author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or
republish, to post on servers or to redistribute to lists, requires prior specific permission the SSL framework. Many previous studies have utilized time as
and/or a fee. Request permissions from permissions@[Link]. an input feature but have ignored its effect on other confounding
ICAIF ’23, November 27–29, 2023, Brooklyn, NY, USA factors.
© 2023 Copyright held by the owner/author(s). Publication rights licensed to ACM.
ACM ISBN 979-8-4007-0240-2/23/11. . . $15.00
Next, in the evolving landscape of investing, traditional strategies
[Link] such as geographical diversification and sector-based investments

533
ICAIF ’23, November 27–29, 2023, Brooklyn, NY, USA Yoontae Hwang, Junhyeong Lee, Daham Kim, Seunghwan Noh, Joohwan Hong, and Yongjae Lee

are becoming increasingly complex due to globalization and digital- In the temporal domain, recent research on the SSL method has
ization, blurring regional and sector boundaries [19]. This complex- predominantly concentrated on video understanding[17] or action
ity is heightened by information asymmetry prevalent in emerging classification[26]. Consequently, there has been limited exploration
markets [1][2][9][10], along with rapid sector alterations spurred by of incorporating periodic information from time-series data. While
technological advancements. Additionally, disruptive technologies TS2VEC[36] has successfully generated robust representations for
such as electric vehicles, blockchains, and AI & robotics are altering individual timestamps through the implementation of a hierarchical
sector dynamics, resulting in alterations to investment instruments, approach with contrastive learning, it is still difficult to incorporate
like thematic exchange-traded funds[5]. Hence, the performance noisy and non-stationary time-series data (e.g., stock data).
of traditional methodologies that rely on traditional regional and
sector classifications should be limited in their effectiveness. In the 2.2 Temporal domain generalization
new era of globalization and digitalization, investment management Domain Generalization(DG) refers to the learning of a general
should go beyond regional and sector classifications and effectively model representation, and various methods have been proposed for
utilize diverse data types. this purpose. Typical methods include data manipulation[30][31],
To achieve good representations of stock data while addressing which involves augmenting input data to acquire a general repre-
these challenges, we propose SimStock. It is based on SSL with sentation; domain-invariant representation learning, such as kernel
temporal domain generalization while being able to utilize various method adversarial training[12][13] and explicit feature alignment
types of data related to stocks. Through this, SimStock is able to between source domains[23] and several learning strategies aimed
efficiently match similar stocks with comprehensive consideration at enhancing generalization power, such as meta-learning[11] and
of their price dynamics as well as other relevant information. We ensemble learning[24]. However, these studies assume that the do-
perform extensive numerical experiments to demonstrate the strong main index set spans time (i.e., temporal) and cannot adaptively
performance of SimStock. learn temporal shifts over time. In this regard, DRAIN[3] is the
We list four key contributions of our paper as follows: first temporal domain generalization (TDG) method to address this
(1) We combine self-supervised learning with temporal domain limitation by adaptively learning temporal drifts across multiple
generalization. Through this, SimStock can identify similar source domains.
stocks for any query stock going beyond temporal, regional, Our proposed model SimStock is based on DRAIN, but there
and sector boundaries. are two main differences between DRAIN and SimStock. First,
(2) We propose a novel corruption method for SSL of stock SimStock incorporates TDG into the SSL framework. Hence, it is
data, termed dimension corruption. By integrating temporal label-free and can take into account both temporal context and
patterns into the corruption process, SimStock can learn static information. Second, we consider the case of adapting from
robust representations of stocks in spite of their noisy and a source domain to multiple target domains (i.e., different stock
non-stationary nature. exchanges) to achieve a universal representation for stock data,
(3) SimStock achieves state-of-the-art performance in finding while DRAIN only considers different temporal domains.
similar stocks across four real-world benchmarks with thou-
sands of stocks. 3 SIMSTOCK
(4) SimStock significantly simplifies the process of screening po- We propose SimStock, which is graphically illustrated in Figure 1. A
tential investment opportunities. For example, SimStock can key distinguishing characteristic of our framework, setting it apart
easily find Chinese stocks that are similar to Nvidia (listed from prior research, is its utilization of stock augmentation and
in NASDAQ, U.S.), or it can find U.S. stocks that are similar temporal domain generalization techniques specifically designed
to Samsung Electronics (listed in KOSPI, South Korea). to capture the dynamic nature of stock data.

2 RELATED WORKS 3.1 Preliminary


We review related works on self-supervised learning and temporal We consider a self-supervised task where the stock data distribution
domain generalization. evolves over time. In the training phase, we are given T observed
source domains 𝐷 1:𝑇 = {D1, D2 , ..., DT }, which are sampled from
2.1 Self-supervised learning distributions at T different time points 𝑡 1 ≤ 𝑡 2 ≤ ... ≤ 𝑡𝑇 . Each
𝑁𝑠
Self-Supervised Learning(SSL) methods have been extensively in- source domain is denoted as Ds = {𝑥𝑖𝑠 , 𝑐𝑖𝑠 }𝑖=1 , for 𝑠 = 1, 2, ..., T ,
𝑠 𝑑
where 𝑥 ∈ R represent the 𝑑𝑚 -dimensional temporal features,
vestigated in the fields of computer vision[7][18][38] and natural 𝑚

language processing[21][27][28][29]. In general, the main objec- 𝑐 𝑠 ∈ R𝑑𝑛 is 𝑑𝑛 -dimensional static metadata, and 𝑁𝑠 is the sample
tive in SSL is to find an embedding space where positive pairs(or size at timestamp 𝑡𝑠 . We have omitted the sample index 𝑖 for sim-
views) of data points remain close to each other, while negative plicity. The model will only be tested on a target domain in the
pairs(or views) are far apart. However, limited research has been future, i.e., DT +1 where 𝑡𝑇 +1 ≥ 𝑡𝑇 .
conducted regarding the application of such techniques to finan- Our goal is to proactively capture the drift from temporal do-
cial data reflecting its temporal characteristics. One of the main mains to find stock representations that are robust with respect to
reasons for this is the challenge of generating different views (both temporal distribution shifts. We presume that the representation
positive and negative), which play a key role in self-supervised model, denoted as 𝑓𝜃𝑠 , is characterized by a deep neural network
representation in non-stationary time-series data (e.g., stock data). with function parameters 𝜃𝑠 at timestamp 𝑡𝑠 . Consequently, we

534
SimStock : Representation Model for Stock Similarities ICAIF ’23, November 27–29, 2023, Brooklyn, NY, USA

Input Temporal Feature Variant Feature Tokenizer Dimension Corruption Representation Triplet Loss Next Domain

Positive view
Open1 Open2 Open𝑘
𝐬
𝐇𝒑𝒐𝒔 𝑓𝜃𝑠 𝑠
𝐂𝐋𝐒𝑃𝑜𝑠
High1 High2 High𝑘 Add
LSTM (𝑔𝜙 )
Low1 Low2 … Low𝑘 𝐇𝒔 Shared weight 𝐇𝒔 (domain s+1)
Close1 Close2 Close𝑘 Negative view
𝐒 𝑠
Price features 𝒙𝒔 Volume1 Volume2 Volume𝑘 TKE 𝑠 =
[𝐶𝐿𝑆] 𝐇𝒏𝒆𝒈 𝑓𝜃𝑠 𝐂𝐋𝐒𝑁𝑒𝑔
𝐓𝐊𝐄𝒊𝒔
Self Supervised Learning
Static Embeddings Decoding Encoding
Noise
Function Function

Static metadata 𝒄𝒔
LSTM (𝑔𝜙 ) LSTM (𝑔𝜙 )

(domain 1) (domain s)
Temporal Domain Generalization

Figure 1: The proposed model (SimStock) combines self-supervised learning framework with temporal domain generalization
for stock representations.

can get the representation embedding 𝑧𝑠 = 𝑓𝜃𝑠 (𝑥 𝑠 , 𝑐 𝑠 ), where 𝑧𝑠 Theorem 1. (Bai et al., [3]) Consider training domains 𝐷 1:𝑇
represent both temporal and static features of stock data. In the where the variance Var(𝐷𝑠 ) is the same for all 𝑠 ∈ {1, ...,𝑇 }. We
next section, we show that the representation model serves as a can establish an inequality for the variance of the predictive dis-
mapping function during training, which predicts the dynamics tribution, which represents each method’s predictive uncertainty.
across the parameter 𝜃 1:𝑇 = {𝜃 1, 𝜃 2, .., 𝜃𝑇 } at each domain 𝐷𝑠 . Specifically, we have: Var(𝑀DRAIN ) < Var(𝑀on ) ≤ Var(𝑀off ). In
this inequality, 𝑀on and 𝑀off denote online and offline learning
3.2 Temporal Domain Generalization models, respectively.
The implication of Theorem 1 extends to the SSL framework.
We are motivated by DRAIN[3], which first proposed the concept
Consequently, we can achieve temporal domain generalization in
of temporal domain generalization. In each temporal domain D𝑠 ,
SSL by finding the optimal model parameters 𝜃𝑇 +1 as described
the representation network 𝑓𝜃𝑠 can be trained by maximizing the
above.
conditional probability P(𝜃𝑠 |D𝑠 ). Here, 𝜃𝑠 signifies the state of the
model parameters at timestamp 𝑡𝑠 . Given the dynamic nature of
D𝑠 , the conditional probability P(𝜃𝑠 |D𝑠 ) will also change over time. 3.3 Temporal Representation Learning
The objective in the context of temporal domain generalization is Our ultimate goal is to learn a representation model, 𝑓𝜃𝑠 , which cap-
to estimate 𝜃𝑇 +1 utilizing all the training data from D1:T . From a tures the stock data distribution that evolves over time. To achieve
probabilistic perspective, we can express this as: this, we develop an SSL framework for temporal representation
learning of stock data.

Temporal feature variant. The time-varying patterns of stock
P(𝜃𝑇 +1 |D1:𝑇 ) = P(𝜃𝑇 +1 |𝜃 1:𝑇 , D1:𝑇 ) · P(𝜃 1:𝑇 |D1:𝑇 )𝑑𝜃 1:𝑇 , (1)
Ω prices are essential for identifying short- and long-term character-
istics of stocks. To learn more rich representations, a price feature
where Ω denotes the space for model parameters 𝜃 1:𝑇 . In Eq. 1, 𝑥 𝑠 is processed by a temporal transformation module 𝜇. Specifi-
the first term inside the integral P(𝜃𝑇 +1 |𝜃 1:𝑇 , D1:𝑇 ) represents the cally, the price feature 𝑥 𝑠 is provided with 𝑘 variations, denoted
inference phase, which is the process of predicting the future state as 𝜇 (𝑥 𝑠 ) = CONCAT(𝜇 1 (𝑥 𝑠 ), 𝜇2 (𝑥 𝑠 ), ..., 𝜇𝑘 (𝑥 𝑠 )) ∈ R𝑑𝑚𝑘 . Here,
of the target representation network (i.e., 𝜃𝑇 +1 ) given all historical 𝑑𝑚𝑘 = 𝑑𝑚 × 𝑘, and each 𝜇 1, 𝜇2, ..., 𝜇𝑘 ∈ U , where U denotes the col-
states (i.e., 𝜃 1:𝑇 , D1:𝑇 ). The second term P(𝜃 1:𝑇 |D1:𝑇 ) signifies the lection of temporal transformations. This module is used to create
training phase, which involves leveraging all training data 𝐷 1:𝑇 to temporal features that incorporate various time intervals. For exam-
ascertain the state of the model on each source domain. ple, 𝜇 1 (𝑥 𝑠 ) and 𝜇2 (𝑥 𝑠 ) would reflect temporal patterns within a day
Suppose that we are at time 𝑡𝑠 . In order to effectively address and a week. Various methods, such as the moving average[33][34],
the temporal drift present across the domain, the next parameters Fourier transform[41], and mixtures of experts[41], can be utilized
𝜃𝑠+1 need to be updated on the current and previous domains 𝐷 1:𝑠 . to create these temporal features. In this study, we use moving
The main problem is how to actually update 𝜃𝑠+1 . In this regard, average, which is the most common choice.
DRAIN introduces a sequential learning process using LSTM[16] to Combined embedding with static metadata. In our frame-
describe the stochastic process of 𝜃𝑠 . Within the LSTM, each unit work, static metadata 𝑐 𝑠 , which can include sectors, company de-
𝑔𝜙 defined by its parameters 𝜙 is used to generate 𝜃𝑠+1 while taking scriptions, and the 3-statement financial data (i.e., income statement,
into consideration the preceding context 𝐷 1:𝑠 and 𝜃 1:𝑠 . This process balance sheet, and cash flow statement), are handled in the static
is illustrated with yellow boxes in Figure 1. embedding layer. As a result, an embedding Embed(𝑐 𝑠 ) ∈ R𝑑𝑚𝑘 is

535
ICAIF ’23, November 27–29, 2023, Brooklyn, NY, USA Yoontae Hwang, Junhyeong Lee, Daham Kim, Seunghwan Noh, Joohwan Hong, and Yongjae Lee

obtained. Different embedding models can be used depending on parameters 𝜃𝑠 through the process described in Section 3.2. In order
the type of data. For example, Ada[25] can incorporate information to effectively reflect temporal patterns of corrupted token embed-
from multiple texts simultaneously, which can be a great tool for dings (H𝑠𝑝𝑜𝑠 and H𝑠𝑛𝑒𝑔 ), we use the self-attention mechanism[32].
practitioners in the finance industry who handle various types of The self-attention mechanism aggregates corrupted token embed-
data. dings with normalized importance as follows:
Next we create a combined embedding that incorporates both
𝑄𝐾𝑇
the temporal feature variant 𝜇 (𝑥 𝑠 ) and the embedded static meta- Attention(𝑄, 𝐾, 𝑉 ) = Softmax( √ )𝑉 . (7)
data Embed(𝑐 𝑠 ). The resulting combined embedding is denoted as 𝑑
follows: Here, 𝑄 = H𝑠∗𝑊𝑄 ∈ R𝑑 ×𝑑𝑘 , 𝐾 = H𝑠∗𝑊𝐾 ∈ R𝑑 ×𝑑𝑘 and 𝑉 = H𝑠∗𝑊𝑉 ∈
H𝑠 = 𝜇 (𝑥 𝑠 ) + Embed(𝑐 𝑠 ) ∈ R𝑑𝑚𝑘 . (2) R𝑑 ×𝑑 𝑣 represent queries, keys, and values, respectively. Note that
Feature Tokenizer module. We draw inspiration from the H𝑠∗ represents token embeddings (either positive or negative), while
tokenizer approach[14], which transforms input features into to- 𝑊𝑄 , 𝑊𝐾 , and 𝑊𝑉 are learnable matrices that share weights between
ken embeddings to obtain more meaningful representations. The positive and negative token embeddings. The output, which has a
feature-wise token embeddings TKE𝑠𝑗 for a given feature index 𝑗 dimension of 𝑑 𝑣 , is then transformed back into an embedding of
are computed as follows: dimension 𝑑 through a fully connected layer. Finally, the outputs
CLS𝑠𝑛𝑒𝑔 and CLS𝑠𝑝𝑜𝑠 are obtained.
TKE𝑠𝑗 = 𝑏𝑠𝑗 + H𝑠𝑗 𝑊 𝑗𝑠 (3)
Triplet loss. For SimStock, we train it to minimize a triplet
where 𝑏𝑠𝑗 ∈ R𝑑 is the 𝑗-th feature bias term and 𝑊 𝑗𝑠 ∈ R𝑑 is the loss[4], which is a popular choice in SSL. The key idea behind
weight vector for the 𝑗-th feature. Consequently, the token embed- triplet loss is the use of triplets, each of which consists of an anchor,
dings TKE𝑠 ∈ R𝑑𝑐 ×𝑑 can be obtained by stacking all of the feature and positive and negative views. Here, the anchor is the embeddings
embeddings and adding a special classification [CLS] token, which for the unperturbed combined embedding.
is known to possess the essence of information after training. This For the triplet (CLS𝑠𝑝𝑜𝑠 , CLS𝑠𝑛𝑒𝑔 , H𝑠 ), where CLS𝑠𝑝𝑜𝑠 is the posi-
is represented as: tive view, CLS𝑠𝑛𝑒𝑔 is the negative view, and H𝑠 is the combined
embedding (anchor), the triplet loss is defined as follows:
TKE𝑠 = STACK([CLS], TKE𝑠1, ..., TKE𝑑𝑠 ) (4)
𝑚𝑘
𝐿triplet = max(0, sim(H𝑠 , CLS𝑠𝑝𝑜𝑠 ) − sim(H𝑠 , CLS𝑠𝑛𝑒𝑔 ) + 𝛼) (8)
where R𝑑𝑐 ×𝑑 = R (𝑑𝑚𝑘 +1) ×𝑑 denotes the dimension of the combined
token embeddings TKE𝑠 . In the above equation, sim(·, ·) denotes a similarity measure (e.g.,
Dimension corruption. When generating views, mixup[39] cosine similarity or Euclidean distance), and 𝛼 > 0 is a margin that
or cutmix[37] methods are most commonly used. These methods is introduced to separate positive pairs from negative pairs. The
are suitable for invariant augmentation of static data (e.g., images), intuition behind this loss function is that we want to ensure that the
however, these are not suitable for time-series data (e.g., stocks). anchor point gets closer to the positive sample than to the negative
We generate views for temporal variants on the same instance, sample by at least the margin 𝛼.
unlike conventional SSL methods that use invariant augmentation Inference phase. In our framework, the inference phase is
by using different instances together. For time-series data, mix- particularly important. Unlike most existing contrastive represen-
ing different sequences would ruin the entire temporal structure. tation learning studies[7][15], our model is specifically designed
Therefore, we propose a dimension corruption method for the aug- to be robust with respect to temporal distribution shifts. The in-
mentation of temporal data. ference phase consists of passing the target domain 𝐷𝑠+1 through
First, we create positive and negative views, H𝑠𝑝𝑜𝑠 and H𝑠𝑛𝑒𝑔 , by the embedding module to obtain the combined embedding H𝑠+1
randomly shuffling the dimensions within the token embeddings and further processed by the feature tokenizer module to obtain
TKE𝑠 . Here, we define two permutation matrices, P𝑠𝑝𝑜𝑠 and P𝑠𝑛𝑒𝑔 , the token embeddings TKE𝑠+1 . The stock representation is then
obtained by feeding TKE𝑠+1 into the representation model 𝑓𝜃𝑠+1 ,
both of size 𝑑 × 𝑑. 1
which is updated with the optimal parameters 𝜃𝑠+1 generated by
the TDG method described in Section 3.2.
H𝑠𝑝𝑜𝑠 = 𝜆TKE𝑠 + (1 − 𝜆)TKE𝑠 P𝑠𝑝𝑜𝑠 (5)
H𝑠𝑛𝑒𝑔 = (1 − 𝜆)TKE𝑠 + 𝜆TKE𝑠 P𝑠𝑛𝑒𝑔 (6) 4 EXPERIMENT
In this case, the formulas (5) and (6) generate positive and nega- Now we present experiment results to thoroughly demonstrate
tive views for self-supervised learning. The degree of this perturba- the performance of SimStock on real-world benchmark datasets.
tion in both views is determined by the mixing parameter 𝜆. With The source code is available at [Link]
𝜆 > 0.5, the positive view H𝑠𝑝𝑜𝑠 has minor perturbations, maintain- SimStock-Representation-Model-for-Stock-Similarities
ing much of the original token embedding. The negative view H𝑠𝑛𝑒𝑔
is more altered, with greater dimension shuffling, deviating more 4.1 Implementation details
from the original. We set 𝜆 = 0.7 as the default value in this study. We present the details of datasets, baseline models, hyperparameter
Representation module. The representation module 𝑓𝜃𝑠 aims selection, training details, evaluation metrics, and the experiment
to characterize the shift between different domains by refining the setting.
1A permutation matrix is a square 0-1 matrix that has exactly one entry of 1 in each Datasets. We collected the daily stock price (OHLCV) and sec-
row and each column and 0s elsewhere. tor information for stocks listed on the NYSE (New York Stock

536
SimStock : Representation Model for Stock Similarities ICAIF ’23, November 27–29, 2023, Brooklyn, NY, USA

US US to SSE US to SZSE US to TSE


4 × 100
3 × 100
101 101 101
2 × 100

100 100 100


100
DTW@10 DTW@5 DTW@3 DTW@1 DTW@10 DTW@5 DTW@3 DTW@1 DTW@10 DTW@5 DTW@3 DTW@1 DTW@10 DTW@5 DTW@3 DTW@1
SSE to US SSE SSE to SZSE SSE to TSE
4 × 100 101 101
101 3 × 100
2 × 100
100 100
100 100
6 × 10 1
DTW@10 DTW@5 DTW@3 DTW@1 DTW@10 DTW@5 DTW@3 DTW@1 DTW@10 DTW@5 DTW@3 DTW@1 DTW@10 DTW@5 DTW@3 DTW@1
SZSE to US SZSE to SSE SZSE SZSE to TSE
101 4 × 100 101
101 3 × 100
2 × 100

100 100 100 100


DTW@10 DTW@5 DTW@3 DTW@1 DTW@10 DTW@5 DTW@3 DTW@1 DTW@10 DTW@5 DTW@3 DTW@1 DTW@10 DTW@5 DTW@3 DTW@1
TSE to US TSE to SSE TSE to SZSE TSE
101 101 101

100
100 100
100
DTW@10 DTW@5 DTW@3 DTW@1 DTW@10 DTW@5 DTW@3 DTW@1 DTW@10 DTW@5 DTW@3 DTW@1 DTW@10 DTW@5 DTW@3 DTW@1
SimStock Corr1 Corr2 Peer TS2VEC

Figure 2: Performance of models in one-to-one (diagonal) and one-to-many (off-diagonal) scenarios for finding similar stocks.
Each data point is accompanied by a 95% confidence interval.

Exchange), NASDAQ (National Association of Securities Dealers February 13, 2018, to February 15, 2022, while the test period spans
Automated Quotations), SSE (Shanghai Stock Exchange), SZSE from February 16, 2022, to May 19, 2023.
(Shenzhen Stock Exchange), and TSE (Tokyo Stock Exchange) from Baseline models. The most widely used method for finding
Yahoo Finance. Table 1 provides detailed information on the pre- similar stocks in financial markets would be to calculate the cor-
processing of price features. relation of stock returns. We use two versions of this correlation
baseline depending on the lookback period. Corr1 uses the past
Price features Description one-year returns, and Corr2 uses returns from the beginning of
𝑧 Open Open𝑡 /Close𝑡 − 1 the test period (i.e., from February 13th, 2018). Another baseline
𝑧 High High𝑡 /Close𝑡 − 1
Peer is the list of similar stocks provided by Google Finance, Yahoo
𝑧 Low Low𝑡 /Close𝑡 − 1
𝑧 Close Close𝑡 /Close𝑡 −1 − 1 Finance, and Financial Modeling Prep. The last baseline model is
𝑧 Volume Volume𝑡 /Volume𝑡 −1 − 1 the state-of-the-art method TS2VEC[36], which is also based on
Table 1: Normalized temporal price features. SSL.
Training detail. All models are trained using the Adam opti-
Here, we generate normalized input features describing the trend mizer [20] with an initial learning rate of 10 −3 . The temporal source
of a stock on day t. 𝑧 Open , 𝑧 High and 𝑧 Low represent the comparison domains are divided on an annual basis, and thus, we have four
values of the opening, highest, and lowest prices, respectively, rela- source domains in the training set.
tive to the closing price of the same day. Also, 𝑧 Close and 𝑧 Volume Evaluation metrics. To assess the performance of our models,
represent the comparative values of the closing prices and the vol- we use Dynamic Time Warping(DTW)[6] to measure the distance
ume values compared with day t-1, respectively. In addition, we between any two stocks. DTW is one of the most widely used
calculated OHLCV for 5, 10, 15, 20, 25, and 30-day intervals for the distance measures for time-series data, because it is more flexible
temporal feature variant. than correlation[40].
We used all stocks listed on the NYSE and NASDAQ (4,231 Consider two time-series sequences 𝑋 = {𝑥 1, ..., 𝑥𝑚 } and 𝑌 =
stocks) and refer to them as the US exchanges in our study. How- {𝑦1, ..., 𝑦𝑛 }. The DTW between 𝑋 and 𝑌 is defined as:
ever, there is one distinguishing feature between the two Chinese
√︄ ∑︁
stock exchanges, SSE (1,407 stocks) and SZSE (1,696 stocks). While 2
𝐷𝑇𝑊 (𝑋, 𝑌 ) = 𝑥𝑖 − 𝑦 𝑗 . (9)
SSE allows access for foreign investors, SZSE does not in general.
(𝑖,𝑗 ) ∈𝜋
Therefore, in our experimental environment, the SSE and SZSE on
the Chinese stock exchange are treated separately. Lastly, we use Here, an alignment path 𝜋 of length 𝐾 is a sequence of 𝐾 index
all stocks listed on TSE (3,882 stocks). The training period is from pairs (𝑖, 𝑗)𝐾 , where max(𝑚, 𝑛) ≤ 𝐾 ≤ 𝑚 + 𝑛 − 1. Also, ||.|| is the

537
ICAIF ’23, November 27–29, 2023, Brooklyn, NY, USA Yoontae Hwang, Junhyeong Lee, Daham Kim, Seunghwan Noh, Joohwan Hong, and Yongjae Lee

Euclidean distance. DTW uses global path constraints while com- On the other hand, in the case of NVIDIA, SimStock and Peer chose
paring two time-series sequences 𝑋 and 𝑌 . That is, the pairs 𝑖 and 𝑗 companies related to semiconductor or image processing. However,
are constrained so that |𝑖 − 𝑗 | ≤ 𝑟 , where 𝑟 is a predefined radius, in TS2VEC, Corr1, and Corr2 methods selected stocks that have no
the case of the Sakoe–Chiba band. Additionally, we evaluate using single relationship nor business similarity to NVDA (such as Houli-
DTW measures by selecting the top 10, 5, 3, and 1 similar stocks han Lokey, CAE Inc, Independence Realty Trust, etc.).
(namely, DTW@10, DTW@5, DTW@3, and DTW@1). Overall, we can see that for a given query stock, SimStock can
find stocks that are similar to the query stock in terms of both
4.2 Can SimStock find similar stocks? time-series distance and fundamental information. Note that all the
In this section, we consider two different scenarios. In the one-to- baseline models are good at only one of them.
one scenario, given a query stock, we find similar stocks within the
same exchange. In the one-to-many scenario, given a query stock, 4.3 Application to index tracking of thematic
we find similar stocks within another exchange. That is, we apply EFTs
the trained weights of a model for one-to-one scenarios to stock Recently, thematic ETFs (e.g., ARK Innovation ETF or Global X
data from another exchange. For example, models trained on the Robotics & AI ETF) have gained popularity among many retail
US exchange can be used to find similar stocks in the SSE, SZSE, or investors who wish to make a bet for some specific investment
TSE exchanges. The query is not restricted to individual stocks. It themes. While there are many different investment themes, the
can be either sector indices or ETFs. Notice that all performances most popular themes were about innovative technologies. Such
are measured in an out-of-sample manner. That is, similar stocks thematic ETFs try to identify innovative tech companies, and thus,
are found based on the training period, but the DTWs are calculated it makes others to track these ETFs.
using returns in the test period. In this section, we compare the performance of SimStock and
4.2.1 One-to-one scenario. The diagonal plots in Figure 2 illustrate Corr2, which showed the best performance among baseline models
the performance (DTW) of different models in one-to-one scenario. in Section 4.2, for identifying stocks to track four popular thematic
It is clear that SimStock stands out as the best performer in the one- EFTs. Our goal is to track thematic ETFs by finding stocks from US,
to-one scenario compared to all other baseline models regardless SSE, SZSE, and TSE. In other words, we use SimStock and Corr2
of exchanges (US, SSE, SZSE, and TSE). to find similar stocks using thematic ETFs as queries.
Note that the performances of all the baseline models were not We used four thematic ETFs: ARK Innovation ETF (ARKK), First
much different. It is interesting that the peer stocks picked by vari- Trust Cloud Computing ETF (SKYY), Global X Robotics & AI ETF
ous investment platforms (Peer) were not quite close to the query (BOTS), and Global X Lithium & Battery Tech ETF (LIT). The in-
stocks in terms of DTW. Also, TS2VEC did not show a significant sample and out-of-sample periods are the same as in Section 4.2.
performance difference compared to Corr1, Corr2, and Peer. This
4.3.1 Tracking error. Based on thematic ETF queries, the two meth-
indicates that TS2VEC is not robust with respect to temporal distri-
ods, SimStock and Corr2, find top 𝑘 similar stocks. Then, we create
bution shifts.
equal-weighted portfolios to track the thematic ETFs. We can write
4.2.2 One-to-many scenario. The off-diagonal plots in Figure 2 rep- the end-of-the-period tracking error TE as
resent the outcomes of identifying for similar stocks in exchanges
different from the exchange of the query stock. Note that Peer is not
v
u
t ∑︁ 𝑛
1
available for this scenario, because most trading platforms do not 𝑇𝐸 = (𝑅 𝐼 − 𝑅 𝑃𝑗 ) 2 (10)
provide information on similar stocks in other exchanges. Again, 𝑛 𝑗=1 𝑗
SimStock consistently outperforms the baseline models. While
TS2Vec or Corr2 perform quite well in some cases (e.g., SZSE to where 𝑅 𝐼𝑗 and 𝑅 𝑃𝑗 are the cumulative return of the query ETF and
SSE, TSE to SSE, SSE to SZSE, SZSE to TSE), they are quite bad in tracking portfolio at period 𝑗, respectively, and 𝑛 is the number of
other cases. periods.
In Table 3, we consider three scenarios. First, choose candidate
4.2.3 Qualitative evaluation. The similarity between stocks should stocks from all exchanges (US, SSE, SZSE, TSE). Second, choose
not be measured only on time-series distances. Similar stocks should stocks from only US exchanges. Finally, choose stocks only from
also have similar fundamental information (such as business area non-US exchanges (SSE, SZSE, TSE). In all three scenarios, SimStock
or 3-statements). Unfortunately, however, it is almost impossible outperforms Corr2 in terms of tracking error for any 𝑘 values (@15,
to measure such similarity in a quantitative way. Therefore, we @20, @25, @30, @35). These results suggest that SimStock can be
provide qualitative analysis of the results. used for index tracking as well.
Table 2 shows the case when different models found the top
5 similar stocks for two query stocks J.P. Morgan and NVIDIA 4.3.2 Qualitative evaluation. Similar to Section 4.2.3, we take a
(NASDAQ: JPM and NVDA). First of all, SimStock and Peer recom- look at the actual stocks that are chosen by SimStock and Corr2.
mended banking stocks, such as Bank Of America, Wells Fargo, and However, unlike Section 4.2.3, the queries are ETFs, not individual
Citi, JP Morgan. Also, Corr1 identified stocks in the financial sector, stocks. Hence, the models should find stocks that are in accordance
such as Gladstone Capital and Metlife, as similar stocks. However, with the theme of the query ETF.
all the top 5 stocks suggested by Corr2 were not from the financial Table 4 presents the top 5 similar stocks found by SimStock and
sector. Corr2 for each Query ETF in the all-exchange case. It is interesting

538
SimStock : Representation Model for Stock Similarities ICAIF ’23, November 27–29, 2023, Brooklyn, NY, USA

Top 𝑘 similar stocks


Query stocks Methods
@1 @2 @3 @4 @5
SimStock Bank of America Wells Fargo & Co Citigroup Deutsche Bank Banco Bilbao
TS2VEC Phillips 66 National Western Life Insurance Celanese Corp Suzano S.A Primo Water
J.P. Morgan
Corr1 First Citizens Bancshares Gladstone Capital Metlife PNC Financial Services Equitable Holdings
(JPM)
Corr2 Ryder System EZCORP Clean Harbors Evercore Lincoln National Corp
Peer Mastercard Incorporated VISA Bank of America Wells Fargo & Co Citigroup
SimStock Applied Materials Axcelis Technologies Advanced Micro Devices Lattice Semiconductor Universal Display Corp
TS2VEC Cinemark Holdings CAE Inc Cimpress KLA Corp Overstock
NVIDIA
Corr1 Houlihan Lokey Synaptics Old Dominion Freight Line Marvell Technology BJ’s Wholesale Club
(NVDA)
Corr2 Skyline Champion Advanced Micro Devices Microsoft Independence Realty Trust MaxLinear
Peer ASML Holding Adobe Cisco Systems Intel Texas Instruments
Table 2: Top 𝑘 similar stocks for query stocks J.P. Morgan and NVIDIA using different methods.

ARK Innovation ETF (ARKK) First Trust Cloud Computing ETF (SKYY)
Exchange Methods Exchange Methods
@15 @20 @25 @30 @35 @15 @20 @25 @30 @35
SimStock 0.0896 0.0861 0.0866 0.0840 0.0900 SimStock 0.0671 0.0667 0.0705 0.0687 0.0703
All-Exchange All-Exchange
Corr2 0.1095 0.1106 0.1262 0.1398 0.1557 Corr2 0.1479 0.1214 0.1314 0.1296 0.1379
SimStock 0.0866 0.0939 0.1069 0.1069 0.1063 SimStock 0.0921 0.1140 0.1124 0.1124 0.1124
US US
Corr2 0.2871 0.2546 0.2331 0.2504 0.2513 Corr2 0.2637 0.1818 0.1912 0.1994 0.1823
SimStock 0.3361 0.3509 0.3597 0.3597 0.3597 SimStock 0.1804 0.1980 0.2020 0.2024 0.1948
Non-US Non-US
Corr2 0.5584 0.5029 0.4931 0.4859 0.4703 Corr2 0.3457 0.3159 3.6853 3.0985 2.6743

Global X Robotics & AI ETF (BOTZ) Global X Lithium & Battery Tech ETF (LIT)
Exchange Methods Exchange Methods
@15 @20 @25 @30 @35 @15 @20 @25 @30 @35
SimStock 0.0512 0.0548 0.0527 0.0530 0.0599 SimStock 0.0603 0.0650 0.0646 0.0636 0.0630
All-Exchange All-Exchange
Corr2 0.0922 0.0958 0.1013 0.1120 0.1067 Corr2 0.1378 0.1073 0.0961 0.1013 0.0899
SimStock 0.0761 0.0842 0.0806 0.0837 0.0836 SimStock 0.0584 0.0543 0.0543 0.0543 0.0543
US US
Corr2 0.1691 0.1552 0.1692 0.1802 0.1674 Corr2 0.1200 0.0897 0.1048 0.1279 0.1080
SimStock 0.1654 0.1647 0.1616 0.1772 0.1784 SimStock 0.0785 0.0916 0.1077 0.1165 0.1234
Non-US Non-US
Corr2 0.3217 0.2958 0.2615 0.2521 0.2479 Corr2 0.1139 0.1377 0.1201 2.8995 2.5279
Table 3: Tracking errors of SimStock and Corr2 for tracking thematic ETFs in various exchanges.

Top 𝑘 similar stocks


Query ETFs Methods
@1 @2 @3 @4 @5
SimStock Stratasys Zuora 3D Systems Corp Transaction Media Networks Azenta
ARKK
Corr2 China Southern Power Grid Technology BlackLine Twist Bioscience Block MercadoLibre
SimStock ServiceNow Guidewire Software Elastic N.V. Splunk Intuit
SKYY
Corr2 China Southern Power Grid Technology Electro Optic Systems Catalent Sea Limited IDEXX Laboratories
SimStock Pci Technology Group Applied Materials Seiko Corporation Zhejiang Taotao Vehicles Teradyne
BOTZ
Corr2 China Southern Power Grid Technology Artisan Partners Asset Management Yeti Holdings Floor & Decor Holdings Cognex Corporation
SimStock Livent Corp Albemarle Corp IQVIA Holdings Inc KKR & Co Trimble Inc
LIT
Corr2 China Southern Power Grid Technology Great Wall Motor Company Limited Sungrow Power Supply BYD Company Limited Albemarle Corporation
Table 4: Top 𝑘 similar stocks for query ETFs ARKK, SKYY, BOTZ and LIT using SimStock and Corr2.

to note that Corr2 identified China Southern Power Grid Technol- domain generalization to self-supervised learning, SimStock could
ogy(Shanghai: [Link]) as the most similar stock for all Thematic learn robust stock representations with respect to temporal shifts.
ETFs. In addition, most stocks selected by Corr2 do not have a di- Our experiments on real-world benchmark datasets demonstrated
rect relationship with the characteristics of the Thematic ETFs. For the effectiveness of SimStock in finding similar stocks with out-of-
example, Corr2 selected companies such as defense (Electro-optic sample evaluations. SimStock outperformed all baseline models,
system), healthcare (Catalent), and pet-related (IDEXX Laborato- including correlation-based methods and the state-of-the-art SSL
ries) companies. On the other hand, for SimStock, all the stocks model. The quantitative and qualitative evaluation results of the
chosen for SKYY are related to cloud computing. Furthermore, framework showed its potential for practical applications in the
SimStock found some actual ETF constituent stocks (all five for field of stock market analysis and investment decision-making.
SKYY and two for LIT).
ACKNOWLEDGMENTS
This work was supported by the National Research Foundation of
5 CONCLUSION Korea (NRF) grant funded by the Korean government (MSIT) (No.
In this study, we propose a novel method called SimStock, a frame- NRF-2022R1I1A4069163).
work for representing the similarity of stocks using self-supervised
learning and temporal domain generalization. We identified two REFERENCES
major challenges for representation learning of stocks: (1) temporal [1] Muhammad Munir Ahmad, Ahmed Imran Hunjra, Faridul Islam, and Qasim
Zureigat. 2021. Does asymmetric information affect firm’s financing decisions?
distributional shifts, and (2) ambiguity of conventional regional and International Journal of Emerging Markets ahead-of-print (2021).
sector classification. [2] Laura Alfaro, Gonzalo Asis, Anusha Chari, and Ugo Panizza. 2017. Lessons
We were able to overcome the second challenge by using the self- unlearned? Corporate debt in emerging markets. Technical Report. National
Bureau of Economic Research.
supervised learning framework that incorporates both temporal [3] Guangji Bai, Chen Ling, and Liang Zhao. 2022. Temporal Domain Generalization
and static features of stock data. Also, by introducing temporal with Drift-Aware Dynamic Neural Networks. arXiv preprint arXiv:2205.10664

539
ICAIF ’23, November 27–29, 2023, Brooklyn, NY, USA Yoontae Hwang, Junhyeong Lee, Daham Kim, Seunghwan Noh, Joohwan Hong, and Yongjae Lee

(2022). [29] Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu. 2020. Mpnet:
[4] Vassileios Balntas, Edgar Riba, Daniel Ponsa, and Krystian Mikolajczyk. 2016. Masked and permuted pre-training for language understanding. Advances in
Learning local feature descriptors with triplets and shallow convolutional neural Neural Information Processing Systems 33 (2020), 16857–16867.
networks.. In Bmvc, Vol. 1. 3. [30] Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, and Pieter
[5] Alka Banerjee, Steven Schoenfeld, and Joy Yang. 2022. Thematic Indexes: Ex- Abbeel. 2017. Domain randomization for transferring deep neural networks
panding Dimensions of Indexing. The Journal of Beta Investment Strategies 13, 1 from simulation to the real world. In 2017 IEEE/RSJ international conference on
(2022), 69–78. intelligent robots and systems (IROS). IEEE, 23–30.
[6] Donald J Berndt and James Clifford. 1994. Using dynamic time warping to find [31] Jonathan Tremblay, Aayush Prakash, David Acuna, Mark Brophy, Varun Jampani,
patterns in time series.. In KDD workshop, Vol. 10. Seattle, WA, USA:, 359–370. Cem Anil, Thang To, Eric Cameracci, Shaad Boochoon, and Stan Birchfield. 2018.
[7] Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A Training deep networks with synthetic data: Bridging the reality gap by domain
simple framework for contrastive learning of visual representations. In Interna- randomization. In Proceedings of the IEEE conference on computer vision and
tional conference on machine learning. PMLR, 1597–1607. pattern recognition workshops. 969–977.
[8] Xinlei Chen and Kaiming He. 2021. Exploring simple siamese representation [32] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones,
learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all
recognition. 15750–15758. you need. Advances in neural information processing systems 30 (2017).
[9] Sisira RN Colombage and Abdel K Halabi. 2012. Asymmetry of information and [33] Gerald Woo, Chenghao Liu, Doyen Sahoo, Akshat Kumar, and Steven Hoi. 2022.
the finance-growth nexus in emerging markets: empirical evidence using panel Etsformer: Exponential smoothing transformers for time-series forecasting. arXiv
VECM analysis. The Journal of Developing Areas (2012), 133–146. preprint arXiv:2202.01381 (2022).
[10] WA de De Wet. 2004. The role of asymmetric information on investments in [34] Haixu Wu, Jiehui Xu, Jianmin Wang, and Mingsheng Long. 2021. Autoformer: De-
emerging markets. Economic Modelling 21, 4 (2004), 621–630. composition transformers with auto-correlation for long-term series forecasting.
[11] Qi Dou, Daniel Coelho de Castro, Konstantinos Kamnitsas, and Ben Glocker. Advances in Neural Information Processing Systems 34 (2021), 22419–22430.
2019. Domain generalization via model-agnostic learning of semantic features. [35] Qiong Wu, Christopher G Brinton, Zheng Zhang, Andrea Pizzoferrato, Zhenming
Advances in Neural Information Processing Systems 32 (2019). Liu, and Mihai Cucuringu. 2021. Equity2vec: End-to-end deep learning framework
[12] Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo for cross-sectional asset pricing. In Proceedings of the Second ACM International
Larochelle, François Laviolette, Mario Marchand, and Victor Lempitsky. 2016. Conference on AI in Finance. 1–9.
Domain-adversarial training of neural networks. The journal of machine learning [36] Zhihan Yue, Yujing Wang, Juanyong Duan, Tianmeng Yang, Congrui Huang,
research 17, 1 (2016), 2096–2030. Yunhai Tong, and Bixiong Xu. 2022. Ts2vec: Towards universal representation of
[13] Rui Gong, Wen Li, Yuhua Chen, and Luc Van Gool. 2019. Dlow: Domain flow time series. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 36.
for adaptation and generalization. In Proceedings of the IEEE/CVF conference on 8980–8987.
computer vision and pattern recognition. 2477–2486. [37] Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and
[14] Yury Gorishniy, Ivan Rubachev, Valentin Khrulkov, and Artem Babenko. 2021. Youngjoon Yoo. 2019. Cutmix: Regularization strategy to train strong classifiers
Revisiting deep learning models for tabular data. Advances in Neural Information with localizable features. In Proceedings of the IEEE/CVF international conference
Processing Systems 34 (2021), 18932–18943. on computer vision. 6023–6032.
[15] Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre [38] Xiaohua Zhai, Avital Oliver, Alexander Kolesnikov, and Lucas Beyer. 2019. S4l:
Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Self-supervised semi-supervised learning. In Proceedings of the IEEE/CVF interna-
Guo, Mohammad Gheshlaghi Azar, et al. 2020. Bootstrap your own latent-a new tional conference on computer vision. 1476–1485.
approach to self-supervised learning. Advances in neural information processing [39] Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz. 2017.
systems 33 (2020), 21271–21284. mixup: Beyond empirical risk minimization. arXiv preprint arXiv:1710.09412
[16] Sepp Hochreiter and Jürgen Schmidhuber. 1997. Long short-term memory. Neural (2017).
computation 9, 8 (1997), 1735–1780. [40] Yichi Zhang, Mihai Cucuringu, Alexander Y Shestopaloff, and Stefan Zohren.
[17] Simon Jenni, Givi Meishvili, and Paolo Favaro. 2020. Video representation learn- 2023. Robust Detection of Lead-Lag Relationships in Lagged Multi-Factor Models.
ing by recognizing temporal transformations. In Computer Vision–ECCV 2020: arXiv preprint arXiv:2305.06704 (2023).
16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part [41] Tian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang, Liang Sun, and Rong Jin.
XXVIII 16. Springer, 425–442. 2022. Fedformer: Frequency enhanced decomposed transformer for long-term
[18] Longlong Jing and Yingli Tian. 2020. Self-supervised visual feature learning series forecasting. In International Conference on Machine Learning. PMLR, 27268–
with deep neural networks: A survey. IEEE transactions on pattern analysis and 27286.
machine intelligence 43, 11 (2020), 4037–4058.
[19] Woo Chang Kim, Yongjae Lee, and Yoon Hak Lee. 2014. Cost of Asset Allocation
in Equity Market: How Much Do Investors Lose Due to Bad Asset Class Design?
The Journal of Portfolio Management 41, 1 (2014), 34–44.
[20] Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Opti-
mization. [Link]
[21] Kenton Lee, Ming-Wei Chang, and Kristina Toutanova. 2019. Latent re-
trieval for weakly supervised open domain question answering. arXiv preprint
arXiv:1906.00300 (2019).
[22] Yongjae Lee, John RJ Thompson, Jang Ho Kim, Woo Chang Kim, and Francesco A
Fabozzi. 2023. An overview of machine learning for asset management. The
Journal of Portfolio Management 49, 9 (2023), 31–63.
[23] Wen Li, Zheng Xu, Dong Xu, Dengxin Dai, and Luc Van Gool. 2017. Domain
generalization and adaptation using low rank exemplar SVMs. IEEE transactions
on pattern analysis and machine intelligence 40, 5 (2017), 1114–1127.
[24] Massimiliano Mancini, Samuel Rota Bulo, Barbara Caputo, and Elisa Ricci. 2018.
Best sources forward: domain generalization through source-specific nets. In 2018
25th IEEE international conference on image processing (ICIP). IEEE, 1353–1357.
[25] Arvind Neelakantan, Tao Xu, Raul Puri, Alec Radford, Jesse Michael Han, Jerry
Tworek, Qiming Yuan, Nikolas Tezak, Jong Wook Kim, Chris Hallacy, et al.
2022. Text and code embeddings by contrastive pre-training. arXiv preprint
arXiv:2201.10005 (2022).
[26] Rui Qian, Tianjian Meng, Boqing Gong, Ming-Hsuan Yang, Huisheng Wang, Serge
Belongie, and Yin Cui. 2021. Spatiotemporal contrastive video representation
learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and
Pattern Recognition. 6964–6974.
[27] Xipeng Qiu, Tianxiang Sun, Yige Xu, Yunfan Shao, Ning Dai, and Xuanjing Huang.
2020. Pre-trained models for natural language processing: A survey. Science
China Technological Sciences 63, 10 (2020), 1872–1897.
[28] Sebastian Ruder and Barbara Plank. 2018. Strong baselines for neural semi-
supervised learning under domain shift. arXiv preprint arXiv:1804.09530 (2018).

540

You might also like