N Using Meta-Learning
N Using Meta-Learning
2901
Authorized licensed use limited to: Wenzhou University. Downloaded on November 05,2024 at 08:22:25 UTC from IEEE Xplore. Restrictions apply.
of each stock 𝑖 is processed based on one-month sliding, as
presented in ② of Fig. 3.
2902
Authorized licensed use limited to: Wenzhou University. Downloaded on November 05,2024 at 08:22:25 UTC from IEEE Xplore. Restrictions apply.
using the following loss function (2). Universal Model
of all stocks of all stocks of all stocks
at month t+2
at month t at month t+1
Update Update
···
Task Learner Task Learner Task Learner the this designed threshold σ𝑛𝑑 .
Figure 5. Meta-testing with individual task-learner.
2903
Authorized licensed use limited to: Wenzhou University. Downloaded on November 05,2024 at 08:22:25 UTC from IEEE Xplore. Restrictions apply.
𝜎𝑑𝑛 10 percent. Additionally, models applying a meta-learning
of 𝑛 was more dispersed, especially during the China-US
𝑝𝑑 framework showed better performance in terms of prediction
trade war of 2018 and the COVID-19 pandemic of 2020. accuracy, balance accuracy, and weighted F1-score. ResNet
Therefore, our proposed slope-detection labeling method can with a meta-learning framework achieved the best “rise”
indeed provide a clearer reflection of stock market trends. prediction accuracy. For comparison, we evaluated the
Furthermore, the experimental results demonstrate that our models’ ability to distinguish between “rise” and “fall” by
proposed labeling method was able to effectively tolerate merging the labels “rise plus” and “rise” into “rise", and the
drastic market changes. labels “fall plus” and “fall” into “fall”. Table VI shows the
B. Parameter Settings of Experiments accuracy of only using these two levels of labels, “rise” and
In our proposed meta-learning framework, we “fall”. The best results were achieved by applying the meta-
incorporated three different classical neural networks, include learning framework with the universal model.
a fully convolutional network (FCN), a residual neural
TABLE V. PREDICTION ACCURACY WITH FOUR-LEVEL LABELS.
network (ResNet), and a temporal convolutional network
Regular Balance Weighted “rise”
(TCN). In addition, individual and universal models were Models
Accuracy Accuracy F1-Score Precision
implemented in our meta-testing process. Abbreviations for Ind-FCN 45.19 % 26.03 % 39.54 % 48.14 %
these model combinations are presented in Table III. Meta-Ind-FCN 48.59 % 28.29 % 44.29 % 51.30 %
Ind-ResNet 41.44 % 25.98 % 38.57 % 47.73 %
Table IV presents the hyperparameters used in our Meta-Ind-ResNet 45.68 % 26.65 % 41.43 % 48.91 %
experiments. All our learning models used the “Adam” Ind-TCN 36.76 % 28.73 % 39.45 % 56.53 %
optimizer and the “CosineAnnealingLR” scheduling module Meta-Ind-TCN 37.04 % 47.71 % 39.14 % 62.88 %
in the training process. In the meta-training phase, we Uni-FCN 59.82 % 40.30 % 57.37 % 63.87 %
adjusted the task-learner for five epochs, and after every five Meta-Uni-FCN 59.33 % 40.32 % 57.02 % 63.92 %
adjustments, we adjusted the meta-learner once, and this Uni-ResNet 59.82 % 40.30 % 57.36 % 64.09 %
Meta-Uni-ResNet 60.09 % 39.97 % 57.50 % 64.22 %
training process was executed for a total of 100 generations. Uni-TCN 62.06 % 36.94 % 57.31 % 64.70 %
In the meta-testing phase, we adjusted the individual model Meta-Uni-TCN 62.14 % 37.73 % 57.81 % 65.06 %
and the universal model in 50 generations, and then selected
TABLE VI. PREDICTION ACCURACY WITH TWO- LEVEL LABELS.
the model hyper-parameters in the last generation.
Regular Balance Weighted
Models
TABLE III. MODEL ABBREVIATIONS. Accuracy Accuracy F1-Score
Models NNs Learning Architectures Ind-FCN 54.38 % 52.01 % 51.37 %
Ind-NNs Individual + NNs Meta-Ind-FCN 58.79 % 57.52 % 58.01 %
FCN Ind-ResNet 52.92 % 51.52 % 51.97 %
Uni-NNs Universal + NNs
ResNet Meta-Ind-ResNet 54.97 % 53.37 % 53.65 %
Meta Ind-NNs Meta learning + individual + NNs
TCN
Meta Uni-NNs Meta learning + universal + NNs Ind-TCN 60.21 % 58.93 % 59.47 %
Meta-Ind-TCN 72.95 % 73.29 % 73.01 %
TABLE IV. PARAMETER SETTINGS. Uni-FCN 76.31 % 75.85 % 76.24 %
Learning Parameter Value settings Meta-Uni-FCN 75.98 % 75.64 % 75.95 %
α, β, γ 1 × 10−4 Uni-ResNet 76.33 % 75.94 % 76.28 %
Steps updating 𝜙 in Meta-Training 100 Meta-Uni-ResNet 76.48 % 76.10 % 76.43 %
Epochs updating 𝜃 in Meta-Training 5 Uni-TCN 77.42 % 76.95 % 77.35 %
Epochs updating 𝜃 for universal model 50 Meta-Uni-TCN 77.66 % 77.25 % 77.60 %
Epochs updating 𝜃 for individual model 50
Epochs updating 𝜃 in Meta Testing 50
Optimizer Adam 2) Investment Profitability Analysis
Scheduler CosineAnnealingLR We also compared various models in terms of investment
profitability. Our stock trading strategy was to buy stocks
C. Experimental Results held in S&P500 when the price trend signal of the stock was
Here, we evaluate the effectiveness of our proposed meta- predicted to be “rise”. Furthermore, these purchased stocks
learning framework in two aspects. First, we present the were sold if other three signals appeared on the next
prediction accuracy of each model. Second, we address forecasting day with the price trend prediction model. In this
investment profitability based on the prediction results as an study, we use the total cumulated return to show the
investment strategy from 2015-01-01 to 2020-12-31. investment profitability according to the prediction signal of
1) Performance Metrics different models and the proposed stock trading strategy.
Additionally, we evenly distribute the funds (Equally
We address the stock trend prediction problem as a four- Weighted Portfolio) to each stock that is predicted to have a
fold classification problem. The four labels used include signal of “rise”. The value of our assets is calculated as the
“rise plus,” “rise,” “fall,” and “fall plus”. To avoid unfair sum of the value of all stocks held which is calculated by
experimental comparisons due to uneven labeling data, we multiplying number of each stock and the stock price on that
applied regular accuracy, balance accuracy, and weighted day. In assessing profitability, we compared the cumulative
F1-score to comprehensively evaluate the performance of return of different models from 2015-01-01 to 2020-12-31.
our proposed models. As presented in Fig. 8, the investment profitability of models
Table V shows that the accuracy of the universal models applying meta-learning frameworks with individual tuning
was better than that of the individual models by more than model achieved better performance than those of the
2904
Authorized licensed use limited to: Wenzhou University. Downloaded on November 05,2024 at 08:22:25 UTC from IEEE Xplore. Restrictions apply.
conventional models. In addition, we have observed that [5] Nikhil Mishra, Mostafa Rohaninejad, Xi Chen, and Pieter Abbeel, “A
when applying a meta-learning framework with universal simple neural attentive meta-learner,” in Proceedings of 6th
International Conference on Learning Representations (ICLR), 2018.
fine tuning models, profitability cannot be significantly [6] Adam Santoro, Sergey Bartunov, Matthew Botvinick, Daan Wierstra,
improved. Different degrees of profit improvement are and Timothy Lillicrap, “Meta-learning with memory-augmented
caused by the accuracy of the meta-learning framework for neural networks,” in Proceedings of International Conference on
trend prediction. Machine Learning (ICML), pp.1842–1850 2016.
[7] Chelsea Finn, Pieter Abbeel, and Sergey Levine, “Model-agnostic
meta-learning for fast adaptation of deep networks,” in Proceedings of
the 34th International Conference on Machine Learning (ICML),
Sydney, PMLR 70, 2017.
[8] Oriol Vinyals, Charles Blundell, Timothy Lillicrap, Koray
Kavukcuoglu, and Daan Wierstra, “Matching net-works for one shot
learning,” Advances in Neural Information Processing Systems, vol.
29, pp.3630–3638, 2016.
[9] Koch, Gregory, “Siamese neural networks for one-shot image
recognition,” in Proceedings of ICML Deep Learning Workshop, 2015.
[10] Antreas Antoniou, Harrison Edwards, and Amos Storkey, “How to
train your MAML,” in Proceedings of the International Conference
on Learning Representations (ICLR), 2019.
[11] Long J, Shelhamer E, and Darrell T, “Fully convolutional networks
for semantic segmentation,” in Proceedings of IEEE International
conference on computer vision and pattern recognition,” pp.3431-
3440, 2015.
Figure 8. Cumulated return from 2015-01-01 to 2020-12-31. [12] Ismail Fawaz Hassan, Forestier G,Weber J, and Idoumghar L,Muller
PA, “Data augmentation using synthetic data for time series
V. CONCLUSION classification with deep residual networks,” in Proceedings of
International Workshop on advanced analytics and learning on
In this study, we have applied meta-learning to solve the temporal data, 2018.
time-series prediction problem for short-term stock price [13] S Bai, JZ Kolter, and V Koltun, “An empirical evaluation of generic
trends. We have proposed a novel slope-detection labeling convolutional and recurrent networks for sequence modeling,” arXiv
method to divide the data into four categories, including preprint arXiv:1803.01271, 2018.
[14] Qiu, M., and Song, Y, “Predicting the direction of stock market index
“rise plus,” “rise,” “fall,” and “fall plus,” which movement using an optimized artificial neural network model,” PLoS
demonstrated the ability to be more representative of stock one Journal, 11(5), 2016.
market trends. We experimentally compared the [15] L. Zhang, C. Aggarwal, and G.-J. Qi, “Stock price prediction via
discovering multi-frequency trading patterns,” in Proceedings of 23rd
effectiveness of different models on the basis of prediction
ACM SIGKDD International Conference on Knowledge Discovery
accuracy and investment profitability. In terms of prediction and Data Mining, pp.2141–2149, 2017.
accuracy, we focused on regular accuracy, balance accuracy, [16] W. Nuij, V. Milea, F. Hogenboom, F. Frasincar, and U. Kaymak, “An
and weighted F1-scores. Our proposed meta-learning automated framework for incorporating news into stock trading
strategies,” IEEE Transactions on Knowledge and Data Engineering,
framework not only achieved higher accuracy than 2(11): pp.823–835, 2014.
conventional methods, but also improves on their accuracy [17] K. A. Althelaya, E. M. El-Alfy and S. Mohammed,”Evaluation of
in avoiding the prediction of opposite trends. In terms of bidirectional LSTM for short-and long-term stock market prediction,”
investment profitability, the Meta-Ind-TCN model was in Proceedings of 9th International Conference on Information and
Communication Systems (ICICS), pp. 151- 156, Irbid, 2018.
superior to the other models. Meta-Ind-TCN model could be [18] Xiao Dan Zhang, Ang Li, and Ran Pan, “Stock Trend Prediction:
trained satisfactorily to achieve greater than 70% accuracy Based on a New Status Method and AdaBoost Probability Support
with two-level labels. Furthermore, the cumulated return of Vector Machine,” Applied Soft Computing, vol. 49, pp.385-398, 2018.
using Meta-Ind-TCN model was able to grow more than 1.8 [19] Jingyi Shen and M. Omair Shafiq, “Short-term stock market price
trend prediction using a comprehensive deep learning system,”
times the original investment amount. Thus, our proposed Journal of Big Data, Article number: 66, 2020.
method demonstrated a significant improvement. In our [20] Xiao Teng, Tuo Wang, Xiang Zhang, Long Lan, Zhigang Luo,
future work, we will seek to improve forecast accuracy and “Enhancing Stock Price Trend Financial via a Time Sensitive Data
Augmentation Method,” Complexity, Article ID 6737951, 2020.
profit performance further by considering other features, [21] Trafalis, T “Short Term Forecasting with Support Vector Machines
such as other technical indices and financial news. and Application to Stock Price Prediction,” in Proceedings of
REFERENCES International Journal of General Systems, 37(6), pp.677-687, 2008.
[1] I. F. Hassan, G. Forestier, J. Weber, L. Idoumghar, and P. A. Muller, [22] Fischer, Thomas, and Christopher Krauss. "Deep learning with long
“Deep learning for time series classification: a review,” Data Mining short-term memory networks for financial market predictions,”
and Knowledge Discovery, vol. 33(4), pp. 917–963, Jul. 2019. European Journal of Operational Research, 270.2: 654-669, 2018.
[2] Chauvet, Marcelle, and Simon Potter, “Coincident and leading [23] Hoseinzade, Ehsan, Saman Haratizadeh, and Arash Khoeini. “U-
indicators of the stock market,” Journal of Empirical Finance, 7.1, cnnpred: A universal cnn-based predictor for stock markets,” arXiv
pp.87-111, 2000. preprint arXiv:1911.12540, 2019.
[3] Yi Cao, Yuhua Li, Sonya Coleman, Ammar Belatreche, and Thomas [24] Lunde, Asger, and Allan Timmermann, “Duration dependence in
Martin McGinnity, “Adaptive hidden Markov model with anomaly stock prices: An analysis of bull and bear markets,” Journal of
states for price manipulation detection,” IEEE Transactions on Neural Business \& Economic Statistics, 22.3, pp.253-273, 2004.
Networks and Learning Systems, 26 (2), pp.318-330, 2015. [25] Pagan, Adrian R., and Kirill A. Sossounov, “A simple framework for
[4] Tsendsuren Munkhdalai and Hong Yu, “Meta networks,” in analysing bull and bear markets,” Journal of Applied Econometrics,
Proceedings of International Conference on Machine Learning 18.1, pp.23-46, 2003.
(ICML), pp.2554-2563, 2017.
2905
Authorized licensed use limited to: Wenzhou University. Downloaded on November 05,2024 at 08:22:25 UTC from IEEE Xplore. Restrictions apply.
Meta-learning enhances feature adaptability through a sliding mechanism in training and fine-tuning processes. By continually updating the meta-learner with new data and adjusting task-learners based on shifted periods, the framework allows models to quickly adapt previous features to align with new data patterns, improving prediction outcomes and relevancy over time .
Meta-learning frameworks improve investment profitability as they enhance prediction accuracy, allowing more precise buy and sell decisions in stock trading, thus achieving better returns. For instance, the Meta-Ind-TCN model demonstrated superior profitability, achieving a cumulated return that was 1.8 times greater than the original investment by effectively utilizing the improved accuracy from the meta-learning framework for trend prediction .
The novel slope-detection labeling method enhances stock market trend prediction by dividing data into four categories: “rise plus,” “rise,” “fall,” and “fall plus,” instead of the traditional binary labels. This method captures more nuanced market movements, allowing the model to better represent stock market trends and potentially improve prediction accuracy and profitability .
The individual model transfers the pre-trained meta-learner to each task-learner specific to each stock, allowing for targeted fine-tuning based on individual stock data. In contrast, the universal model applies the pre-trained meta-learner to a single task-learner, with all training sets used to fine-tune this task-learner, aiming for a generalized approach applicable to any stock. This distinction allows the individual model to be more specialized, while the universal model strives for broader applicability .
The use of four-level labels, such as “rise plus” and “fall plus,” enhances the model's prediction accuracy by providing a more granular representation of market conditions, reducing the chances of misclassification between subtle rises and falls. This method improves the model's ability to differentiate between varying magnitudes of trend movements, resulting in higher prediction accuracy .
A two-dimensional input tensor is used in the model to integrate 11 technical indices and OHLC prices over 22 days for each stock, forming a comprehensive view of past trends to predict future stock movements. This structured integration of data allows for efficient processing and model training, enhancing the prediction accuracy by considering a wide range of input parameters in the analysis .
During the meta-training phase, the query set is used to evaluate the trained task-learners. The summation of the loss from each task-learner when the query set is input is calculated and applied to update the meta-learner's parameters. This process ensures the meta-learner adapts effectively based on real data outcomes .
The framework supports real-time predictions by pre-training the initial parameters of each task-learner based on prior knowledge, assuming similar patterns will recur close to the target period. This pre-training leverages historical data patterns, ensuring task-learners are immediately more effective when adapting to real-time data, thus optimizing prediction outcomes continuously .
To prevent overfitting during task-learner training, the meta-learning framework keeps the number of epochs low. This strategy limits excessive adjustments of the task-learner’s parameters to the training data, which helps maintain generalization ability and avoids fitting noise or specific patterns irrelevant to new data .
The meta-learning framework addresses dataset imbalance by applying a sampled support set in the first phase of the meta-training process to avoid training failure. The support set is organized by sampling an equal number of records in each labeling category to ensure balanced data input. This approach influences the adaptability of the update process of the meta-learner, improving the model's capability to generalize to imbalanced datasets .