מדידת חפיפה בין צינורות נתונים של משגר הטוקנים בסולנה
סיכום
המחקר בוחן אם שני אוספי נתונים שהוגדרו בנפרד צופים באותם טוקנים במשגר pump.fun של Solana. צינור אחד מזהה קבוצות הקשורות לארנקי קונים מוקדמים שפועלים בתיאום; השני מתעד החלטות של מסנן דחייה לפני עסקה מצד הסוחר. המחברים משווים קבוצות מזהים של טוקנים בשני חלונות עוקבים, מגבילים את רשומות הדחייה לחלון של כל קבוצה ומיישמים ארבעה כללי חותמת זמן. הם מדווחים על ספירות, שיעורים שנצפו ומדדי חפיפה של Jaccard ו-Dice.
התוצאות מראות חפיפה מועטה מאוד, אך הכיסוי בחלונות שונה מאוד. בחלון השני, 5 מתוך 623 טוקנים שזוהו בקבוצה מופיעים גם בין 1,742 טוקני הדחייה. בראשון, פער באיסוף הנתונים מותיר כיסוי של 14.25 השעות האחרונות בלבד; טוקן אחד חופף בין הקבוצות המלאות, ואילו בתוך פרק הזמן המכוסה אין חפיפה. הכיסוי הקצר הופך את ההשוואה הראשונה לחסרת מידע. לכן המחקר מזהיר שהגדרות האוסף מעצבות את מדגם הטוקנים שנצפה, וממליץ לדווח על הכיסוי לצד החפיפה בין אוספים. ההשוואות נוגעות לצינורות הנתונים ולחלונות האלה, ולא לכל אוספי הנתונים של Solana.
רעיונות מרכזיים
- המחקר משווה קבוצות מזהי טוקנים מגלאי קבוצות ומתיעוד דחיות לפני עסקה.
- הצינורות מציגים חפיפה מוגבלת מאוד בשני חלונות התצפית.
- פער באיסוף הנתונים הופך את ההשוואה בתוך פרק הכיסוי של החלון הראשון לקטנה מכדי לפרשה.
- הגדרות האוסף עשויות להשפיע על הטוקנים הכלולים במדגמי מחקר בשרשרת.
- המחברים ממליצים לדווח על כיסוי האוסף ועל החפיפה בין אוספים.
תגיות
הטקסט המלא
# Do Two On-Chain Observation Pipelines See the Same Tokens? Cross-Pipeline Coverage on the Solana pump.fun Launchpad # Do Two On-Chain Observation Pipelines See the Same Tokens? Cross-Pipeline Coverage on the Solana pump.fun Launchpad On-chain studies of memecoin launchpads usually rely on one data-collection pipeline, yet whether differently configured pipelines observe the same tokens is rarely measured. This paper compares the output mint sets of two separately configured pipelines from one research programme on the Solana pump.fun launchpad: a cohort-detection pipeline that flags coordinated early-buyer wallets, and a rejection-filtering pipeline that logs a trader-side observer's pre-trade filter decisions. Two consecutive, non-overlapping windows are analysed (v1: June 2026; v2: June-July 2026); in each, the rejection data are restricted to the interval spanned by the cohort detections under four timestamp rules. Only raw counts, observed proportions and Jaccard/Dice indices are reported. In v2, where the rejection stream covers the whole 17.5-day window, the cohort set contains 623 mints and the rejection set 1,742, with 5 overlapping mints (0.803% of cohort; 0.287% of rejection). In v1 a collection gap means the rejection stream covers only the final 14.25 hours (4.4%) of the 13.4-day cohort window; the 20,162 cohort and 53 rejection mints share 1 mint, and within the covered interval none, a count too small to be informative. Both windows show nearly disjoint outputs, agreeing in direction but not in magnitude. Collector configuration, not only market behaviour, therefore shapes which tokens an on-chain study observes, and the paper proposes reporting cross-collector overlap and collector coverage as a routine check. Data and a script that re-derives every number are openly available.
מוצג במלואו בציון המקור ובהתאם לרישיון שלו. רישיון: abstract CC0
הסיכום נכתב בידי סוכן המחקר של Stratmill על סמך המקור; הוא אינו העתק של המקור.