Skip to content
All library documents

Measuring Token Overlap Between Solana Launchpad Data Pipelines

Article arXiv papers · Author: Arati Uday Kamat

Summary

This study asks whether two independently configured data collectors observe the same tokens on the Solana pump.fun launchpad. One pipeline detects cohorts associated with coordinated early-buyer wallets; the other records decisions from a trader-side pre-trade rejection filter. The authors compare mint sets over two consecutive windows, restricting the rejection records to each cohort window and applying four timestamp rules. They report counts, observed proportions, and Jaccard and Dice overlap measures.

The results show very little overlap, but the windows differ sharply in coverage. In the second window, 5 of 623 cohort mints also appear among 1,742 rejection mints. In the first, a collection gap leaves only the final 14.25 hours covered; one mint overlaps across the full sets, while none overlap within the covered interval. That short coverage makes the first comparison uninformative. The study therefore cautions that collector setup shapes observed token samples, and recommends reporting coverage alongside cross-collector overlap. Its comparisons concern these pipelines and windows, not all Solana data collectors.

Key ideas

  • The study compares token mint sets from a cohort detector and a pre-trade rejection logger.
  • The pipelines show very limited overlap in both observation windows.
  • A collection gap makes the first window’s within-coverage comparison too small to interpret.
  • Collector configuration can shape the tokens included in on-chain research samples.
  • The authors recommend reporting collector coverage and cross-collector overlap.

Tags

Full text
# Do Two On-Chain Observation Pipelines See the Same Tokens? Cross-Pipeline Coverage on the Solana pump.fun Launchpad


# Do Two On-Chain Observation Pipelines See the Same Tokens? Cross-Pipeline Coverage on the Solana pump.fun Launchpad









On-chain studies of memecoin launchpads usually rely on one data-collection pipeline, yet whether differently configured pipelines observe the same tokens is rarely measured. This paper compares the output mint sets of two separately configured pipelines from one research programme on the Solana pump.fun launchpad: a cohort-detection pipeline that flags coordinated early-buyer wallets, and a rejection-filtering pipeline that logs a trader-side observer's pre-trade filter decisions. Two consecutive, non-overlapping windows are analysed (v1: June 2026; v2: June-July 2026); in each, the rejection data are restricted to the interval spanned by the cohort detections under four timestamp rules. Only raw counts, observed proportions and Jaccard/Dice indices are reported. In v2, where the rejection stream covers the whole 17.5-day window, the cohort set contains 623 mints and the rejection set 1,742, with 5 overlapping mints (0.803% of cohort; 0.287% of rejection). In v1 a collection gap means the rejection stream covers only the final 14.25 hours (4.4%) of the 13.4-day cohort window; the 20,162 cohort and 53 rejection mints share 1 mint, and within the covered interval none, a count too small to be informative. Both windows show nearly disjoint outputs, agreeing in direction but not in magnitude. Collector configuration, not only market behaviour, therefore shapes which tokens an on-chain study observes, and the paper proposes reporting cross-collector overlap and collector coverage as a routine check. Data and a script that re-derives every number are openly available.

Shown in full with attribution under the source's licence. Licence: abstract CC0

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.