נושאי חדשות תלויי־הקשר לגורמי תמחור נכסים
סיכום
מחקר זה בודק אם הקשר ברמת המשפט משפר מודלים של נושאי חדשות המשמשים לבניית גורמי תמחור נכסים שיטתיים. הוא משווה הקצאת דיריכלה לטנטית (LDA) למודל Sentence Transformer שקובעו משקולותיו ולאחריו אשכולות k-means, תוך שימוש באותו אוסף של 394,661 כתבות ובאותו תהליך לבניית תיק בהמשך. הגישות נבדלות גם בהיקף הטקסט מהכתבות שהן משתמשות בו ובאופן דירוג מונחי הנושא.
גישת הממיר רשמה קוהרנטיות גבוהה יותר של הנושאים, שנמדדה באמצעות NPMI, וציוני שארפ גבוהים יותר לתיק, אך הבדיקות הזמינות אינן מוכיחות שהיא גוברת על LDA. גרסאות ניסיוניות עם אשכולות כדוריים וחשיפות על פני אופקים מרובים הניבו מודל משולב עם שארפ תשואה עודפת של 1.03. הממצאים מרמזים שייצוגים המודעים להקשר עשויים לסייע בחילוץ אותות שימושיים פיננסית מחדשות. רמת הביטחון נותרת מוגבלת: המחקר קורא לבדיקות מחמירות יותר, המגבילות את הקלט למידע שהיה זמין בכל תאריך, ולהערכה על מערכי נתונים רחבים יותר.
רעיונות מרכזיים
- המחקר משווה נושאים שהופקו באמצעות LDA לנושאים שהופקו מהטמעות של Sentence Transformer וקובצו באמצעות k-means.
- בניתוח המדווח, גם קוהרנטיות הנושאים וגם ציוני שארפ של התיק היו גבוהים יותר בגישת הממיר.
- ההשוואה משתמשת באותו אוסף כתבות ובאותו תהליך בניית תיק בהמשך, אך קלט הטקסט ודירוג המונחים שונים.
- אשכולות כדוריים ניסיוניים וחשיפות במספר אופקים הניבו שארפ תשואה עודפת של 1.03 למודל המשולב.
- הראיות אינן מכריעות, ונדרשות בדיקות המבוססות על מידע שהיה זמין בזמן אמת ועל מערכי נתונים רחבים יותר.
תגיות
הטקסט המלא
# From Word Counts to Context: Topic Models for Asset Pricing # From Word Counts to Context: Topic Models for Asset Pricing News may reveal systematic risk, but whether its context enhances the construction of systematic risk factors is still unclear. We seek to test whether utilizing a sentence transformer represents an improvement over techniques such as Latent Dirichlet Allocation (LDA) in the coherence of topic term lists generated from unstructured text data. To test this, the same collection of unstructured text data comprising of 394,661 articles and the same downstream financial portfolio construction pipeline were applied with the text layer differing, including the length of article text each model used and how topic terms were ranked: we benchmark LDA against a frozen sentence transformer with k-means clustering. We find that the sentence transformer branch had higher observed scores both in terms of coherence (measured by NPMI) as well as financial performance (measured by Sharpe), although the available tests do not establish outperformance. Further exploratory specifications such as utilizing spherical clustering and multi-horizon exposures had an observed excess-return Sharpe of 1.03 for the combined model. We believe that there is some promise in applying context-aware techniques on unstructured news text, but stricter tests using only information available at each date and broader datasets may be required to enhance the confidence in the observed performance.
מוצג במלואו בציון המקור ובהתאם לרישיון שלו. רישיון: abstract CC0
הסיכום נכתב בידי סוכן המחקר של Stratmill על סמך המקור; הוא אינו העתק של המקור.