Dataset Preview
Duplicate
The full dataset viewer is not available (click to read why). Only showing a preview of the rows.
The dataset generation failed
Error code:   DatasetGenerationError
Exception:    ArrowInvalid
Message:      Failed to parse string: '物/化/生(3选1)' as a scalar of type double
Traceback:    Traceback (most recent call last):
                File "/usr/local/lib/python3.14/site-packages/datasets/builder.py", line 1837, in _prepare_split_single
                  writer.write_table(table)
                  ~~~~~~~~~~~~~~~~~~^^^^^^^
                File "/usr/local/lib/python3.14/site-packages/datasets/arrow_writer.py", line 765, in write_table
                  self._write_table(pa_table, writer_batch_size=writer_batch_size)
                  ~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                File "/usr/local/lib/python3.14/site-packages/datasets/arrow_writer.py", line 773, in _write_table
                  pa_table = table_cast(pa_table, self._schema)
                File "/usr/local/lib/python3.14/site-packages/datasets/table.py", line 2369, in table_cast
                  return cast_table_to_schema(table, schema)
                File "/usr/local/lib/python3.14/site-packages/datasets/table.py", line 2303, in cast_table_to_schema
                  cast_array_to_feature(
                  ~~~~~~~~~~~~~~~~~~~~~^
                      table[name] if name in table_column_names else pa.array([None] * len(table), type=schema.field(name).type),
                      ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                      feature,
                      ^^^^^^^^
                  )
                  ^
                File "/usr/local/lib/python3.14/site-packages/datasets/table.py", line 1852, in wrapper
                  return pa.chunked_array([func(chunk, *args, **kwargs) for chunk in array.chunks])
                                           ~~~~^^^^^^^^^^^^^^^^^^^^^^^^
                File "/usr/local/lib/python3.14/site-packages/datasets/table.py", line 2143, in cast_array_to_feature
                  return array_cast(
                      array,
                  ...<2 lines>...
                      allow_decimal_to_str=allow_decimal_to_str,
                  )
                File "/usr/local/lib/python3.14/site-packages/datasets/table.py", line 1854, in wrapper
                  return func(array, *args, **kwargs)
                File "/usr/local/lib/python3.14/site-packages/datasets/table.py", line 2006, in array_cast
                  return array.cast(pa_type)
                         ~~~~~~~~~~^^^^^^^^^
                File "pyarrow/array.pxi", line 1147, in pyarrow.lib.Array.cast
                File "/usr/local/lib/python3.14/site-packages/pyarrow/compute.py", line 412, in cast
                  return call_function("cast", [arr], options, memory_pool)
                File "pyarrow/_compute.pyx", line 604, in pyarrow._compute.call_function
                File "pyarrow/_compute.pyx", line 399, in pyarrow._compute.Function.call
                  result = GetResultValue(
                File "pyarrow/error.pxi", line 155, in pyarrow.lib.pyarrow_internal_check_status
                  return check_status(status)
                File "pyarrow/error.pxi", line 92, in pyarrow.lib.check_status
                  raise convert_status(status)
              pyarrow.lib.ArrowInvalid: Failed to parse string: '物/化/生(3选1)' as a scalar of type double
              
              The above exception was the direct cause of the following exception:
              
              Traceback (most recent call last):
                File "/src/services/worker/src/worker/job_runners/config/parquet_and_info.py", line 1369, in compute_config_parquet_and_info_response
                  parquet_operations, partial, estimated_dataset_info = stream_convert_to_parquet(
                                                                        ~~~~~~~~~~~~~~~~~~~~~~~~~^
                      builder, max_dataset_size_bytes=max_dataset_size_bytes
                      ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                  )
                  ^
                File "/src/services/worker/src/worker/job_runners/config/parquet_and_info.py", line 948, in stream_convert_to_parquet
                  builder._prepare_split(split_generator=splits_generators[split], file_format="parquet")
                  ~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                File "/usr/local/lib/python3.14/site-packages/datasets/builder.py", line 1683, in _prepare_split
                  for job_id, done, content in self._prepare_split_single(
                                               ~~~~~~~~~~~~~~~~~~~~~~~~~~^
                      gen_kwargs=gen_kwargs, job_id=job_id, **_prepare_split_args
                      ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                  ):
                  ^
                File "/usr/local/lib/python3.14/site-packages/datasets/builder.py", line 1869, in _prepare_split_single
                  raise DatasetGenerationError("An error occurred while generating the dataset") from e
              datasets.exceptions.DatasetGenerationError: An error occurred while generating the dataset

Need help to make the dataset viewer work? Make sure to review how to configure the dataset viewer, and open a discussion for direct support.

province
string
province.1
string
year
int64
batch
string
category
string
university_code
null
university_name
string
major_group
null
major_code
float64
major_name
string
major_note
null
subject_req
null
plan_count
int64
duration
string
tuition
float64
安徽
安徽
2,017
本科一批
理科
null
北京大学
null
0
工科试验班类
null
null
1
-
0
安徽
安徽
2,017
本科一批
理科
null
北京大学
null
0
理科试验班类
null
null
1
-
0
安徽
安徽
2,017
本科一批
理科
null
北京大学
null
201
经济学类
null
null
1
-
0
安徽
安徽
2,017
本科一批
理科
null
北京大学
null
null
法学
null
null
1
-
0
安徽
安徽
2,017
本科一批
理科
null
北京大学
null
701
数学类
null
null
1
-
0
安徽
安徽
2,017
本科一批
理科
null
北京大学
null
702
物理学类
null
null
1
-
0
安徽
安徽
2,017
本科一批
理科
null
北京大学
null
703
化学类
null
null
1
-
0
安徽
安徽
2,017
本科一批
理科
null
北京大学
null
708
地球物理学类
null
null
1
-
0
安徽
安徽
2,017
本科一批
理科
null
北京大学
null
807
电子信息类
null
null
2
-
0
安徽
安徽
2,017
本科一批
理科
null
北京大学
null
825
环境科学与工程类
null
null
1
-
0
安徽
安徽
2,017
本科一批
理科
null
北京大学
null
1,202
工商管理类
null
null
2
-
0
安徽
安徽
2,017
本科一批
理科
null
北京大学
null
1,204
公共管理类
null
null
1
-
0
安徽
安徽
2,017
本科一批
理科
null
北京大学
null
30,202
国际政治
null
null
1
-
0
安徽
安徽
2,017
本科一批
理科
null
北京大学
null
70,401
天文学
null
null
1
-
0
安徽
安徽
2,017
本科一批
理科
null
北京大学
null
71,001
生物科学
null
null
1
-
0
安徽
安徽
2,017
本科一批
理科
null
北京大学
null
71,101
心理学
null
null
1
-
0
安徽
安徽
2,017
本科一批
理科
null
北京大学
null
82,802
城乡规划
null
null
2
-
0
安徽
安徽
2,017
本科一批
理科
null
北京大学
null
120,102
信息管理与信息系统
null
null
1
-
0
安徽
安徽
2,017
本科一批
文科
null
北京大学
null
0
文科试验班类
null
null
1
-
0
安徽
安徽
2,017
本科一批
文科
null
北京大学
null
101
哲学类
null
null
1
-
0
安徽
安徽
2,017
本科一批
文科
null
北京大学
null
201
经济学类
null
null
1
-
0
安徽
安徽
2,017
本科一批
文科
null
北京大学
null
null
法学
null
null
1
-
0
安徽
安徽
2,017
本科一批
文科
null
北京大学
null
501
中国语言文学类
null
null
2
-
0
安徽
安徽
2,017
本科一批
文科
null
北京大学
null
601
历史学类
null
null
1
-
0
安徽
安徽
2,017
本科一批
文科
null
北京大学
null
1,202
工商管理类
null
null
1
-
0
安徽
安徽
2,017
本科一批
文科
null
北京大学
null
1,204
公共管理类
null
null
1
-
0
安徽
安徽
2,017
本科一批
文科
null
北京大学
null
30,202
国际政治
null
null
2
-
0
安徽
安徽
2,017
本科一批
文科
null
北京大学
null
30,301
社会学
null
null
1
-
0
安徽
安徽
2,017
本科一批
文科
null
北京大学
null
50,201
英语
null
null
1
-
0
安徽
安徽
2,017
本科一批
文科
null
北京大学
null
60,103
考古学
null
null
1
-
0
安徽
安徽
2,017
本科一批
文科
null
北京大学
null
120,102
信息管理与信息系统
null
null
1
-
0
安徽
安徽
2,017
本科一批
理科
null
中国人民大学
null
0
理科试验班(信息与数学)
null
null
5
-
0
安徽
安徽
2,017
本科一批
理科
null
中国人民大学
null
0
理科试验班(环境理工、环境经管)
null
null
3
-
0
安徽
安徽
2,017
本科一批
理科
null
中国人民大学
null
0
理科试验班(物理、化学与心理学)
null
null
2
-
0
安徽
安徽
2,017
本科一批
理科
null
中国人民大学
null
0
理科试验班
null
null
3
-
0
安徽
安徽
2,017
本科一批
理科
null
中国人民大学
null
0
理科试验班(2)
null
null
2
-
0
安徽
安徽
2,017
本科一批
理科
null
中国人民大学
null
0
理科试验班(3)
null
null
5
-
0
安徽
安徽
2,017
本科一批
理科
null
中国人民大学
null
201
经济学类
null
null
3
-
0
安徽
安徽
2,017
本科一批
理科
null
中国人民大学
null
203
金融学类
null
null
3
-
0
安徽
安徽
2,017
本科一批
理科
null
中国人民大学
null
203
金融学类(2)
null
null
5
-
0
安徽
安徽
2,017
本科一批
理科
null
中国人民大学
null
null
法学
null
null
2
-
0
安徽
安徽
2,017
本科一批
理科
null
中国人民大学
null
null
政治学、经济学与哲学
null
null
1
-
0
安徽
安徽
2,017
本科一批
理科
null
中国人民大学
null
503
新闻传播学类
null
null
2
-
0
安徽
安徽
2,017
本科一批
理科
null
中国人民大学
null
712
统计学类
null
null
3
-
0
安徽
安徽
2,017
本科一批
理科
null
中国人民大学
null
1,204
公共管理类
null
null
2
-
0
安徽
安徽
2,017
本科一批
理科
null
中国人民大学
null
30,202
国际政治
null
null
2
-
0
安徽
安徽
2,017
本科一批
理科
null
中国人民大学
null
null
政治学、经济学与哲学(PPE实验班)
null
null
1
-
0
安徽
安徽
2,017
本科一批
理科
null
中国人民大学
null
1,202
工商管理类
null
null
4
-
0
安徽
安徽
2,017
本科一批
理科
null
中国人民大学
null
120,206
人力资源管理
null
null
2
-
0
安徽
安徽
2,017
本科一批
理科
null
中国人民大学
null
120,503
信息资源管理
null
null
2
-
0
安徽
安徽
2,017
本科一批
文科
null
中国人民大学
null
0
人文科学试验班
null
null
3
-
0
安徽
安徽
2,017
本科一批
文科
null
中国人民大学
null
201
经济学类
null
null
4
-
0
安徽
安徽
2,017
本科一批
文科
null
中国人民大学
null
203
金融学类
null
null
3
-
0
安徽
安徽
2,017
本科一批
文科
null
中国人民大学
null
null
法学
null
null
2
-
0
安徽
安徽
2,017
本科一批
文科
null
中国人民大学
null
303
社会学类
null
null
2
-
0
安徽
安徽
2,017
本科一批
文科
null
中国人民大学
null
503
新闻传播学类
null
null
3
-
0
安徽
安徽
2,017
本科一批
文科
null
中国人民大学
null
1,202
工商管理类
null
null
5
-
0
安徽
安徽
2,017
本科一批
文科
null
中国人民大学
null
30,202
国际政治
null
null
3
-
0
安徽
安徽
2,017
本科一批
文科
null
中国人民大学
null
null
政治学、经济学与哲学
null
null
1
-
0
安徽
安徽
2,017
本科一批
文科
null
中国人民大学
null
null
政治学、经济学与哲学(PPE实验班)
null
null
1
-
0
安徽
安徽
2,017
本科一批
文科
null
中国人民大学
null
1,203
农业经济管理类
null
null
3
-
0
安徽
安徽
2,017
本科一批
文科
null
中国人民大学
null
1,204
公共管理类
null
null
3
-
0
安徽
安徽
2,017
本科一批
文科
null
中国人民大学
null
50,201
英语
null
null
2
-
0
安徽
安徽
2,017
本科一批
文科
null
中国人民大学
null
120,206
人力资源管理
null
null
2
-
0
安徽
安徽
2,017
本科一批
文科
null
中国人民大学
null
120,503
信息资源管理
null
null
2
-
0
安徽
安徽
2,017
本科一批
理科
null
清华大学
null
0
理科试验班类(新雅书院)
null
null
1
-
0
安徽
安徽
2,017
本科一批
理科
null
清华大学
null
0
工科试验班类(环境、化工与新材料)
null
null
4
-
0
安徽
安徽
2,017
本科一批
理科
null
清华大学
null
0
工科试验班类(机械、航空与动力)
null
null
9
-
0
安徽
安徽
2,017
本科一批
理科
null
清华大学
null
0
工科试验班类(能源)
null
null
4
-
0
安徽
安徽
2,017
本科一批
理科
null
清华大学
null
0
工科试验班类(自动化与工业工程)
null
null
3
-
0
安徽
安徽
2,017
本科一批
理科
null
清华大学
null
0
理科试验班类(化生)
null
null
2
-
0
安徽
安徽
2,017
本科一批
理科
null
清华大学
null
0
理科试验班类(经济、金融与管理)
null
null
2
-
0
安徽
安徽
2,017
本科一批
理科
null
清华大学
null
0
理科试验班类(人文与社会)
null
null
1
-
0
安徽
安徽
2,017
本科一批
理科
null
清华大学
null
0
理科试验班类(数理)
null
null
4
-
0
安徽
安徽
2,017
本科一批
理科
null
清华大学
null
0
工科试验班类
null
null
9
-
0
安徽
安徽
2,017
本科一批
理科
null
清华大学
null
0
理科试验班类
null
null
1
-
0
安徽
安徽
2,017
本科一批
理科
null
清华大学
null
0
工科试验班类(2)
null
null
4
-
0
安徽
安徽
2,017
本科一批
理科
null
清华大学
null
0
理科试验班类(2)
null
null
4
-
0
安徽
安徽
2,017
本科一批
理科
null
清华大学
null
0
工科试验班类(3)
null
null
3
-
0
安徽
安徽
2,017
本科一批
理科
null
清华大学
null
null
法学
null
null
3
-
0
安徽
安徽
2,017
本科一批
理科
null
清华大学
null
807
电子信息类
null
null
6
-
0
安徽
安徽
2,017
本科一批
理科
null
清华大学
null
809
计算机类
null
null
3
-
0
安徽
安徽
2,017
本科一批
理科
null
清华大学
null
810
土木类
null
null
5
-
0
安徽
安徽
2,017
本科一批
理科
null
清华大学
null
828
建筑类
null
null
2
-
0
安徽
安徽
2,017
本科一批
理科
null
清华大学
null
null
临床医学
null
null
2
-
0
安徽
安徽
2,017
本科一批
理科
null
清华大学
null
null
临床医学(协和)
null
null
2
-
0
安徽
安徽
2,017
本科一批
文科
null
清华大学
null
0
文科试验班类(经济、金融与管理)
null
null
3
-
0
安徽
安徽
2,017
本科一批
理科
null
北京交通大学
null
0
经济管理试验班
null
null
4
-
0
安徽
安徽
2,017
本科一批
理科
null
北京交通大学
null
0
理科试验班类
null
null
12
-
0
安徽
安徽
2,017
本科一批
理科
null
北京交通大学
null
0
文科试验班类
null
null
2
-
0
安徽
安徽
2,017
本科一批
理科
null
北京交通大学
null
null
法学
null
null
1
-
0
安徽
安徽
2,017
本科一批
理科
null
北京交通大学
null
802
机械类
null
null
12
-
0
安徽
安徽
2,017
本科一批
理科
null
北京交通大学
null
802
机械类(2)
null
null
5
-
0
安徽
安徽
2,017
本科一批
理科
null
北京交通大学
null
804
材料类
null
null
5
-
0
安徽
安徽
2,017
本科一批
理科
null
北京交通大学
null
806
电气类
null
null
4
-
0
安徽
安徽
2,017
本科一批
理科
null
北京交通大学
null
807
电子信息类
null
null
12
-
0
安徽
安徽
2,017
本科一批
理科
null
北京交通大学
null
807
电子信息类(通信与控制)
null
null
12
-
0
安徽
安徽
2,017
本科一批
理科
null
北京交通大学
null
807
电子信息类(2)
null
null
2
-
0
安徽
安徽
2,017
本科一批
理科
null
北京交通大学
null
809
计算机类
null
null
7
-
0
安徽
安徽
2,017
本科一批
理科
null
北京交通大学
null
809
计算机类(2)
null
null
2
-
0
End of preview.

English 中文

China College Admission Dataset GitHub Hugging Face Blog RedNote License


English

GaokaoCompass — China College Admission Dataset

GaokaoCompass is a structured dataset of China's national college entrance examination (Gaokao) admission records, covering all 31 provinces from 2017 to 2025. It includes enrollment plans, university admission cutoff scores, major-level admission scores, and score-ranking tables. The dataset is designed to help students, parents, and researchers make informed decisions with transparent, queryable admission data.


Overview

Metric Value
Provinces 31
Year span 2017–2025
Total CSV files 861
Total data rows 11,327,563
Full coverage (all 4 tables) 31 provinces × 2022–2025

Per-table statistics

Table Files Rows Description
enrollment-plan 227 5,862,063 University enrollment plans by province
major-admission 218 4,497,048 Major-level admission scores
school-admission 227 763,558 University-level admission cutoff scores
score-range 189 204,894 Score-ranking tables (candidate distribution)

Directory Structure

data/
├── 2017/
│   ├── anhui/
│   │   ├── enrollment-plan.csv        # Enrollment plan
│   │   ├── school-admission.csv       # University admission cutoff
│   │   ├── major-admission.csv        # Major-level admission scores
│   │   ├── score-range.csv            # Score-ranking table
│   │   └── meta.json                  # Summary statistics (6 dashboard cards)
│   ├── beijing/
│   │   └── ...
│   └── ... (31 provinces)
├── 2018/
│   └── ...
└── 2025/
    └── ...

Table Schemas

1. score-range.csv — Score-Ranking Table

Shows how many candidates scored at each point level, enabling rank lookup.

Field Type Description
province string Province name
year int Year
category string Subject track (理科/文科/物理类/历史类/综合)
batch string Admission batch
control_score int Batch cutoff score
score int Score point
score_range string Score range label
segment_count int Number of candidates at this score
cumulative_count int Cumulative count (rank)
rank_range string Rank range label

2. enrollment-plan.csv — Enrollment Plan

Each row is one university-major enrollment slot in a province.

Field Type Description
province string Province
year int Year
batch string Batch (本科一批/本科二批/专科批 etc.)
category string Subject track
university_code string University code
university_name string University name
major_group string Major group
major_code string Major code
major_name string Major name
major_note string Major notes
subject_req string Subject requirements
plan_count int Planned enrollment count
duration string Program duration
tuition float Annual tuition (CNY)

3. school-admission.csv — University Admission Cutoff

Each row is one university's minimum admission score in a province.

Field Type Description
province string Province
year int Year
category string Subject track
batch string Batch
university_code string University code
university_name string University name
major_group string Major group
subject_req string Subject requirements
min_score int Minimum admission score
min_rank int Minimum admission rank
control_score int Provincial batch cutoff
score_diff int Score above batch cutoff
admit_count int Number admitted
school_province string Province where university is located
school_nature string Ownership (公办 public / 民办 private)
is_985 int Project 985 university (1/0)
is_211 int Project 211 university (1/0)

4. major-admission.csv — Major-Level Admission Scores

Each row is one university-major's admission detail in a province.

Field Type Description
province string Province
year int Year
category string Subject track
batch string Batch
university_code string University code
university_name string University name
major_group string Major group
major_code string Major code
major_name string Major name
major_note string Major notes
subject_req string Subject requirements
min_score int Minimum score
min_rank int Minimum rank
max_score int Maximum score
avg_score float Average score
admit_count int Number admitted
school_province string University location
school_nature string Ownership
is_985 int Project 985 (1/0)
is_211 int Project 211 (1/0)

meta.json — Summary Statistics

Each province/year directory contains a meta.json with pre-computed statistics for 6 dashboard cards.

Card Key fields Description
exam_overview total_candidates, categories, batch_lines Total candidates, batch cutoff scores
university_stats total, is_985, is_211, public, private University counts by type
enrollment_stats total_plan, university_count, top_universities Enrollment plan summary
top_majors name, plan_count, university_count Most popular majors
score_distribution score, segment_count, cumulative_count Sampled score distribution
admission_bands min_score, max_score, median_score Score ranges per batch

Example (Zhejiang 2024):

{
  "exam_overview": {
    "categories": [
      {
        "name": "综合",
        "total_candidates": 281057,
        "max_score": 699,
        "batch_lines": [
          {"batch": "平行录取一段", "score": 492},
          {"batch": "平行录取二段", "score": 269}
        ]
      }
    ],
    "total_candidates": 281057
  },
  "university_stats": {
    "total": 1608,
    "is_985": 37,
    "is_211": 108,
    "public": 900,
    "private": 400
  }
}

Quick Start

Load from Hugging Face

from datasets import load_dataset

dataset = load_dataset("choucsan/Gaokao-Compass-11M")

Read a CSV directly

import pandas as pd

# Load 2024 Zhejiang enrollment plan
df = pd.read_csv("data/2024/zhejiang/enrollment-plan.csv")

# Search for Computer Science majors
cs = df[df["major_name"].str.contains("计算机科学与技术", na=False)]
print(cs[["university_name", "major_name", "plan_count", "tuition"]])

Look up a score ranking

import pandas as pd

# Load 2024 Henan score-ranking table
df = pd.read_csv("data/2024/henan/score-range.csv")

# Find rank for a score of 600
row = df[df["score"] == 600]
print(f"Rank at 600: {row['cumulative_count'].values[0]}")

Use meta.json

import json

with open("data/2024/zhejiang/meta.json", encoding="utf-8") as f:
    meta = json.load(f)

print(f"Total candidates: {meta['exam_overview']['total_candidates']}")
print(f"Universities: {meta['university_stats']['total']}")
print(f"Project 985 universities: {meta['university_stats']['is_985']}")

Data Pipeline

Raw data is sourced from provincial education examination authorities and the National Education Examination Authority (阳光高考). The ETL pipeline:

  1. Discovery — Scan raw Excel files per province, classify by type (enrollment plan / school admission / major admission / score range)
  2. Column normalization — Map 30+ Chinese column name variants to unified English fields
  3. Cleaning — Normalize subject tracks, safe numeric casting, strip decimals from code fields
  4. Deduplication — MD5-based file dedup; row-level dedup by (university, batch, category, major_code)
  5. Year extraction — Parse year from filenames and directory paths
  6. Validation — Null-rate checks, error summary, completeness report

2025 Coverage

Provinces enrollment-plan school-admission major-admission score-range
28 provinces
Qinghai (青海)
Shanxi (山西)
Tibet (西藏)

Use Cases

  • College application — Query eligible universities and majors by score and rank
  • Admission trend analysis — Visualize cutoff score and enrollment changes over years
  • University comparison — Compare admission difficulty across 985/211/public/private institutions
  • Major popularity — Rank majors by enrollment quota and number of offering universities
  • Education research — Academic studies and policy analysis on Gaokao data

中文

高考录取数据平台 · GaokaoCompass

GaokaoCompass 是一个面向中国高考(普通高等学校招生全国统一考试)的结构化录取数据集,覆盖 31 个省份、2017–2025 年的招生计划、院校投档线、专业录取分数和一分一段表数据。旨在为考生、家长和教育研究者提供透明、可查询的高考录取信息参考。


数据概览

指标 数值
省份 31
年份跨度 2017–2025
数据文件总数 861 个 CSV
数据总行数 11,327,563
四表齐全省份/年份 2022–2025 年 31 省全覆盖

各表统计

数据表 文件数 数据行数 说明
enrollment-plan 227 5,862,063 各省高校招生计划
major-admission 218 4,497,048 各高校专业录取分数
school-admission 227 763,558 各高校院校投档线
score-range 189 204,894 一分一段表(考生排名)

数据结构

data/
├── 2017/
│   ├── anhui/
│   │   ├── enrollment-plan.csv        # 招生计划
│   │   ├── school-admission.csv       # 院校投档线
│   │   ├── major-admission.csv        # 专业录取分数
│   │   ├── score-range.csv            # 一分一段表
│   │   └── meta.json                  # 统计摘要(6类卡片)
│   ├── beijing/
│   │   └── ...
│   └── ...(31省)
├── 2018/
│   └── ...
└── 2025/
    └── ...

四类数据表说明

1. score-range.csv · 一分一段表

考生分数排名数据,反映每个分数段的考生人数分布。

字段 类型 说明
province string 省份
year int 年份
category string 科类(理科/文科/物理类/历史类/综合)
batch string 批次
control_score int 批次控制线
score int 分数
score_range string 分数区间
segment_count int 本段人数
cumulative_count int 累计人数(排名)
rank_range string 排名区间

2. enrollment-plan.csv · 招生计划

各高校在各省的招生专业和计划人数。

字段 类型 说明
province string 省份
year int 年份
batch string 批次(本科一批/本科二批/专科批等)
category string 科类
university_code string 院校代码
university_name string 院校名称
major_group string 专业组
major_code string 专业代码
major_name string 专业名称
major_note string 专业备注
subject_req string 选科要求
plan_count int 计划人数
duration string 学制
tuition float 学费(元/年)

3. school-admission.csv · 院校投档线

各高校在各省的最低录取分数和位次。

字段 类型 说明
province string 省份
year int 年份
category string 科类
batch string 批次
university_code string 院校代码
university_name string 院校名称
major_group string 专业组
subject_req string 选科要求
min_score int 最低分
min_rank int 最低位次
control_score int 省控线
score_diff int 批次线差
admit_count int 录取人数
school_province string 学校所在省份
school_nature string 办学性质(公办/民办)
is_985 int 是否 985 高校(1/0)
is_211 int 是否 211 高校(1/0)

4. major-admission.csv · 专业录取分数

各高校各专业在各省的录取分数详情。

字段 类型 说明
province string 省份
year int 年份
category string 科类
batch string 批次
university_code string 院校代码
university_name string 院校名称
major_group string 专业组
major_code string 专业代码
major_name string 专业名称
major_note string 专业备注
subject_req string 选科要求
min_score int 最低分
min_rank int 最低位次
max_score int 最高分
avg_score float 平均分
admit_count int 录取人数
school_province string 学校所在省份
school_nature string 办学性质
is_985 int 是否 985
is_211 int 是否 211

meta.json · 统计摘要

每个省份/年份目录下包含一个 meta.json 文件,提供 6 类卡片的统计信息,适合直接用于前端可视化展示。

卡片 字段 说明
exam_overview total_candidates, categories, batch_lines 高考人数、批次控制线
university_stats total, is_985, is_211, public, private 院校统计
enrollment_stats total_plan, university_count, top_universities 招生计划统计
top_majors name, plan_count, university_count 热门专业排名
score_distribution score, segment_count, cumulative_count 分数分布采样
admission_bands min_score, max_score, median_score 各批次录取分数段

示例(2024 浙江):

{
  "exam_overview": {
    "categories": [
      {
        "name": "综合",
        "total_candidates": 281057,
        "max_score": 699,
        "batch_lines": [
          {"batch": "平行录取一段", "score": 492},
          {"batch": "平行录取二段", "score": 269}
        ]
      }
    ],
    "total_candidates": 281057
  },
  "university_stats": {
    "total": 1608,
    "is_985": 37,
    "is_211": 108,
    "public": 900,
    "private": 400
  }
}

快速使用

从 Hugging Face 加载

from datasets import load_dataset

dataset = load_dataset("choucsan/Gaokao-Compass-11M")

直接读取 CSV

import pandas as pd

# 读取 2024 年浙江的招生计划
df = pd.read_csv("data/2024/zhejiang/enrollment-plan.csv")

# 查询计算机科学与技术专业的招生计划
cs = df[df["major_name"].str.contains("计算机科学与技术", na=False)]
print(cs[["university_name", "major_name", "plan_count", "tuition"]])

查询一分一段表

import pandas as pd

# 读取 2024 年河南理科一分一段表
df = pd.read_csv("data/2024/henan/score-range.csv")

# 查看 600 分对应的排名
row = df[df["score"] == 600]
print(f"600分排名: {row['cumulative_count'].values[0]}")

使用 meta.json

import json

with open("data/2024/zhejiang/meta.json", encoding="utf-8") as f:
    meta = json.load(f)

print(f"浙江2024高考人数: {meta['exam_overview']['total_candidates']}")
print(f"招生院校数: {meta['university_stats']['total']}")
print(f"985高校: {meta['university_stats']['is_985']} 所")

数据来源

数据来自各省教育考试院、阳光高考平台等公开渠道,经清洗、标准化和去重后整理为统一格式。

处理流程:

  1. 文件发现:扫描各省原始 Excel 文件,按类型自动分类
  2. 列名标准化:30+ 种原始列名映射为统一英文字段
  3. 数据清洗:科类标准化、数值安全转换、代码字段去小数点
  4. 去重:按文件大小 + MD5 去除重复文件,专业条目按 (院校, 批次, 科类, 专业代码) 去重
  5. 年份提取:从文件名和目录路径自动提取年份
  6. 质量验证:空值率检查、错误汇总、数据完整性报告

2025 年数据覆盖

省份 enrollment-plan school-admission major-admission score-range
28 省
青海
山西
西藏

应用场景

  • 考生志愿填报:根据分数和排名查询可报考的院校和专业
  • 录取趋势分析:历年分数线、招生人数变化趋势可视化
  • 院校对比:985/211/公办/民办院校在各省的录取难度对比
  • 专业热度分析:各专业招生计划人数和开设院校数量统计
  • 教育研究:高考数据的学术研究和政策分析

许可证

MIT License


联系方式

如有问题、纠错或合作需求:

choucisan@gmail.com

Downloads last month
3,973