wendaostudy/cn-geo-citation-dataset — explained in plain English
Analysis updated 2026-05-18
Study how Chinese generative search platforms select and expose citation sources across different question types.
Analyze citation patterns by research dimension, such as industry, prompt style, or time sensitivity.
Verify data integrity using the provided checksums and manifest before running large-scale analysis.
Compare source exposure behavior across 12 different platform and device combinations.
| wendaostudy/cn-geo-citation-dataset | 23k65a1408/create-aeronautics-skywards | 8015238355/mm2-analytics-dashboard-2026 | |
|---|---|---|---|
| Stars | 185 | 185 | 185 |
| Setup difficulty | easy | moderate | moderate |
| Complexity | 2/5 | 3/5 | 2/5 |
| Audience | researcher | general | general |
Figures from each repo's GitHub metadata at analysis time.
No installation needed beyond a JSON parser, just download and read the JSONL files.
This repository is a research dataset, not a piece of software. It contains 214,119 citation records collected from four major Chinese language model platforms, covering 12 different platform versions across web and mobile apps. Each record captures what a generative search platform cited or retrieved in response to a benchmark question, which makes the dataset useful for studying how these Chinese AI search systems choose their sources and which sources they tend to expose to users. The data is organized into JSON Lines files, split by seven top-level research dimensions such as edge and real-world scenarios, question intent, industry, prompt style, and time sensitivity, with 32 dimension and subcategory pairs in total. Each category is broken into shard files capped at 5,000 records apiece, and every record follows a documented schema with a stable identifier and a hash for verifying data integrity. A machine-readable manifest lists every file, its record count, size, and checksum. The repository includes supporting documentation: a data card describing what the dataset is and how it was built, a data dictionary defining every field, and a quality report covering missing values, duplicate records, and known limitations. Duplicate entries are kept on purpose, to preserve the original shape of the research population rather than artificially cleaning it. A short Python example shows how to iterate through all the record files and parse each line as JSON. This dataset is intended for researchers studying AI search behavior, information retrieval, or citation patterns, rather than for developers looking for a library to install. It is released under the Creative Commons Attribution 4.0 license, though any quoted titles, URLs, or excerpts inside the records may still carry their original publishers' rights, so users need to check applicable copyright and research ethics rules themselves.
A structured dataset of 214,119 citation and retrieval records from four Chinese AI search platforms, for studying how they select and expose sources.
Requires giving credit to the original creator, the underlying quoted content may carry separate rights.
Setup difficulty is rated easy, with roughly 5min to a first successful run.
Mainly researcher.
This repo across BitVibe Labs
double-check against the repo, no cap.