الانتقال إلى المحتوى
جميع مستندات المكتبة

موضوعات الأخبار المراعية للسياق لعوامل تسعير الأصول

مقال arXiv papers · المؤلف: Kevin Foley et al.

الملخص

تختبر هذه الدراسة ما إذا كان السياق على مستوى الجملة يحسّن نماذج موضوعات الأخبار المستخدمة لبناء عوامل تسعير الأصول المنهجية. وتقارن بين تخصيص ديريشليه الكامن ومحوّل جمل ثابت يتبعه تجميع كي-مينز، باستخدام المجموعة نفسها المؤلفة من 394,661 مقالة وعملية بناء المحافظ اللاحقة نفسها. كما تختلف الطريقتان في مقدار نص المقالات المستخدم وكيفية ترتيب مصطلحات الموضوعات.

سجل نهج المحوّل اتساقًا أعلى للموضوعات، مقاسًا بـNPMI، ودرجات شارب أعلى للمحافظ، لكن الاختبارات المتاحة لا تثبت تفوقه على LDA. وأنتجت النسخ الاستكشافية التي تستخدم التجميع الكروي والتعرضات عبر آفاق متعددة نموذجًا مركبًا بلغ شارب العائد الزائد فيه 1.03. وتشير النتائج إلى أن التمثيلات الواعية بالسياق قد تساعد على استخلاص إشارات مفيدة ماليًا من الأخبار. لكن مستوى الثقة يظل محدودًا: إذ تدعو الدراسة إلى اختبارات أشد تقصر المدخلات على المعلومات المتاحة في كل تاريخ، وإلى تقييم مجموعات بيانات أوسع.

الأفكار الرئيسية

  • تقارن الدراسة موضوعات LDA بموضوعات مستخلصة من تضمينات محوّل الجمل المجمعة باستخدام كي-مينز.
  • كان كل من اتساق الموضوعات ودرجات شارب للمحافظ أعلى لنهج المحوّل في التحليل المعلن.
  • تستخدم المقارنة مجموعة المقالات نفسها وخط معالجة بناء المحافظ اللاحق نفسه، رغم اختلاف نص الإدخال وترتيب المصطلحات.
  • أنتج التجميع الكروي الاستكشافي والتعرضات متعددة الآفاق شاربًا للعائد الزائد قدره 1.03 للنموذج المركب.
  • الأدلة غير حاسمة، ويلزم إجراء اختبارات تراعي المعلومات المتاحة في كل تاريخ واستخدام مجموعات بيانات أوسع.

الوسوم

النص الكامل
# From Word Counts to Context: Topic Models for Asset Pricing


# From Word Counts to Context: Topic Models for Asset Pricing









News may reveal systematic risk, but whether its context enhances the construction of systematic risk factors is still unclear. We seek to test whether utilizing a sentence transformer represents an improvement over techniques such as Latent Dirichlet Allocation (LDA) in the coherence of topic term lists generated from unstructured text data. To test this, the same collection of unstructured text data comprising of 394,661 articles and the same downstream financial portfolio construction pipeline were applied with the text layer differing, including the length of article text each model used and how topic terms were ranked: we benchmark LDA against a frozen sentence transformer with k-means clustering. We find that the sentence transformer branch had higher observed scores both in terms of coherence (measured by NPMI) as well as financial performance (measured by Sharpe), although the available tests do not establish outperformance. Further exploratory specifications such as utilizing spherical clustering and multi-horizon exposures had an observed excess-return Sharpe of 1.03 for the combined model. We believe that there is some promise in applying context-aware techniques on unstructured news text, but stricter tests using only information available at each date and broader datasets may be required to enhance the confidence in the observed performance.

يُعرض النص كاملًا مع نسبه إلى مصدره وفقًا لترخيصه. الترخيص: abstract CC0

أعدّ وكيل الأبحاث في Stratmill هذا الملخص استنادًا إلى المصدر الأصلي؛ وهو ليس نسخة منه.