Datasets
Packaged historical extracts for research, backtesting and academic work. The data exists; what does not exist yet is the legal clearance to redistribute it.
No dataset is available for download today. There are no sample files, no CSV or Parquet exports and no commercial dataset licences on offer. This page documents the candidates and, more usefully, the specific reason each one is not published yet.
Why this is not just an engineering task
Building an export is a day of work. Being allowed to distribute what comes out of it is the actual problem, and it is not one that can be solved by shipping anyway.
The layer's inputs arrive under commercial agreements. Those agreements permit Moonlytics to process the data and to build applications on it. They do not automatically permit redistribution of the data — or of extracts close enough to it to substitute for a licence with the original provider. That distinction runs right through the four candidates below.
- Provider terms. Redistribution rights are per-agreement and per-field, and some fields are more constrained than others.
- Derived versus raw. Output we compute ourselves sits on firmer ground than a repackaged provider column. Where the line falls needs confirming per dataset, not assuming.
- News and social content rights. Article text and post content belong to publishers and platforms. The layer holds neither, which is deliberate and which keeps the derived products cleaner.
- Platform terms. Community and channel metrics come with platform terms of their own.
Candidate datasets
Each of these is a real series in the production store today, described with the coverage it actually has.
Hourly asset-level sentiment with the social volume it was computed from.
Blocker: Derived measure over licensed inputs — redistribution scope under review.
Social volume, interactions and dominance as an aligned hourly panel.
Blocker: Closest to raw provider output. Most constrained of the four.
Every retained hourly topic snapshot: topic, source count, category, summary.
Blocker: Our own clustering output over third-party news. Most likely candidate to ship first.
Daily channel, repository and contributor series per project.
Blocker: Collected from public sources; terms per source need checking individually.
What a dataset page will contain
When a dataset does ship, its page will carry all of the following, because a dataset without them is not usable for serious work:
- Description — what a row represents.
- Coverage — assets included and how they were selected.
- Time range — first and last observation, and known gaps.
- Frequency — native grain, with no resampling.
- Schema — every column, its type, its unit, its nullability.
- Format — CSV, JSON or Parquet, with encoding stated.
- Update frequency — whether it is a static extract or refreshed.
- Methodology — how derived columns were computed.
- Licence — what you may and may not do with it, in plain terms.
- Limitations — survivorship, coverage boundaries, known artefacts.
- Sample — a real excerpt, not a schema mock-up.
Attribution
Any dataset released under a licence permitting reuse will ask for attribution: Data by Moonlytics, linking to moonlytics.io. Visible, honest attribution — no hidden links, no keyword-stuffed anchors, no requirement to place it below the fold.
If you need historical data now
Tell us what you actually need — which series, which range, which assets, and what for. A concrete request can be assessed against the relevant provider terms on its merits, which is far more likely to produce something than waiting for a general-purpose release.