J and Q/Kdb+ for Financial Time-Series Data
Summary
The discussion compares J with Q/Kdb+ as tools for quantitative finance, focusing on the database capabilities that make Q/Kdb+ common in trading firms. The accepted answer says J can provide array-oriented computation, but users would need suitable storage and query facilities to match the full workflow offered by Q/Kdb+.
It highlights partitioning as a key feature for large time-series datasets: date-based partitions can limit scans, and parallel processing can speed queries across dates. The answer characterizes JDB as lacking partitioning at the time of the discussion, while noting J’s open-source status and stronger academic presence. A separate reply mentions a commercial J database, and another suggests R with a partitioned database as an alternative. These are practitioner observations, not a benchmark; feature comparisons may also change with software versions and implementations.
Key ideas
- Q/Kdb+ is often used for time-series storage even by firms that make limited use of its language.
- Matching its workflow in J requires database storage plus query and join functionality.
- Date partitioning can make large tick-history queries more selective and practical.
- The answer identifies J’s open-source status and academic use as advantages, alongside lower industry adoption.
- The comparisons are informal and do not provide performance benchmarks.
Tags
Full text
# Can the J language be used as an effective alternative to Q/Kdb+? # Can the J language be used as an effective alternative to Q/Kdb+? I hear a lot about Q/kdb+. I've never had the opportunity to use it for anything real but have played with it using their trial license and found it intriguing (if not somewhat mind warping). I've seen a few references to the J language which is also a derivative of APL according to Wikipedia and have heard of at least one quant trader using it as a poor man's replacement for Q/kdb+. Has anyone put J to use in the real world? Is it worth investing time and energy into to learn, especially if I'm not interested in investing the dollars required to get a kdb+ license? I suspect a lot of K/Q's value comes from the kdb+ database that underlies it. I'm not sure there is an equivalent for J, but I haven't performed an exhaustive search. Interestingly enough, J recently went fully open source and is now hosted (not officially) on Github: https://github.com/openj/core. EDIT: Assuming some folks here are familiar with J and Q/kdb+ but have not used J extensively in industry, what might be some major advantages or disadvantages of using J. EDIT 2: J has a database solution in JDB. One obvious follow-up which I also asked in a comment to @chrisaycock's answer below is how JDB might compare to kdb+ in terms of functionality. ## Answer by chrisaycock (score 7, accepted) https://quant.stackexchange.com/a/1871 You are right that the database component is the main selling point of q/kdb+. Indeed, many Kx customers use kdb+ just for time-series data (and often just for tick history) and don't take advantage of the q language. So the first disadvantage to using J as-is would be to miss the reason most folks even bother with Kx: data storage. Now of course you could clone the data storage for J, just as you could for any language. The really hard part is coming-up with your own version of the query and join routines found in `q.k`, that cryptic file that comes with the q software distribution. Regarding JDB, the biggest difference I can find is that it doesn't support partitioning. For example, kdb+ can split the column-mapped files into separate directories for each date. That means a query `select from quotes where date=2011.09.08` will cause kdb+ to jump directly to today's ticks. Taking that further, kdb+ supports threading to some degree (via `par.txt`) so that it can process multiple dates in parallel. And beyond the performance benefits, partitioning is practically required for really large tables anyway since the columns are memory-mapped at query time; there's no feasible way to do that for a multi-terabyte database without some selectivity. The second disadvantage is that there aren't nearly as many finance shops that use J. Almost every big bank has a Kx license somewhere, so there's always an expert if needed. J doesn't have nearly that adoption. Of course, J's main advantage now is that it's open source. And even before going open, J Software was more academic-friendly than Kx Systems, so there's more presence of J in the universities. Because of that, there seems to be more J examples in papers and on Project Euler than for either k or q. ## Answer by pascalJ (score 1) https://quant.stackexchange.com/a/27580 to add to chrisaycock's reply, J has Jd database (comercial) now. Its a purer vector language than q, but can be used with the same simplicity. ## Answer by Erik Aronesty (score -1) https://quant.stackexchange.com/a/17902 If you're looking for an alternative to Q, try R. It's a a great language for modeling and time series analysis. For a DB, try cassandra, which supports partitioned data, high-speed access/queries. Nice thing is that you can quickly produce compelling graphs and charts in R. Also, R has a big community compared to J. And it's all open source, with a strong, dynamic base of users...so you won't get stuck with code that no one can support or understand.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.