修复README,增加说明
This commit is contained in:
67
README.md
67
README.md
@@ -7,23 +7,48 @@
|
||||
```
|
||||
project/
|
||||
├── data/
|
||||
│ ├── raw_data.json # 原始输入数据(包含噪声和错误)
|
||||
│ └── cleaned_data.json # 清洗后的输出数据
|
||||
│ ├── raw_data_sample.json # 原始输入数据(包含噪声和错误)
|
||||
│ └── cleaned_data_final.json # 清洗后的输出数据
|
||||
├── src/
|
||||
│ └── cleaner.py # 清洗运行程序
|
||||
├── report/
|
||||
│ └── cleaning_report.json # 详细清洗报告
|
||||
│ └── detailed_cleaning_report.json # 详细清洗报告
|
||||
└── README.md # 项目说明文档
|
||||
```
|
||||
|
||||
## 功能特性
|
||||
* **去噪处理**:使用卡尔曼滤波算法平滑连续数值(如通信范围),消除测量噪声。
|
||||
* **异常检测与修复**:
|
||||
* 自动修正负数、极端值。
|
||||
* 强制将效能指标(如打击精度、机动性)限制在 [0, 1] 范围内。
|
||||
* 修复无效的地理坐标。
|
||||
* **缺失值填充**:基于同类型(Role)单位的统计均值进行智能插值。
|
||||
* **标准化**:统一文本格式、时间格式及数值精度。
|
||||
## 清洗流程说明
|
||||
|
||||
程序按以下顺序依次执行五步清洗操作:
|
||||
|
||||
### 1. 去重 (clean_duplicates)
|
||||
- **依据**:按 `TARGET_ID` 去重,同一目标只保留一条记录
|
||||
- **策略**:若存在 `CREATED_TIME` 字段,按创建时间倒序排序后保留最新一条;否则保留第一条
|
||||
- **目的**:消除重复录入的目标节点
|
||||
|
||||
### 2. 缺失值填充 (handle_missing_values)
|
||||
- **范围**:对所有数值型列(除 ID 外)进行处理
|
||||
- **策略**:
|
||||
- 先按 `ROLE_ID` 分组,用同类型单位的均值填充该组内缺失值
|
||||
- 若分组后仍有缺失(如该组仅有一条记录),用全列均值兜底填充
|
||||
- **目的**:保证数值字段完整,便于后续建模
|
||||
|
||||
### 3. 异常值纠正 (correct_outliers)
|
||||
- **涉及字段**:`TARGET_RECOGNITION_CAPABILITY`、`STRIKE_ACCURACY`、`ANTI_JAMMING_CAPABILITY`、`ENVIRONMENT_ADAPTABILITY`、`MOBILITY`
|
||||
- **规则**:
|
||||
- 若值 < 0:取绝对值
|
||||
- 若值 > 1:截断为 1.0
|
||||
- **目的**:将效能指标统一限制在 [0, 1] 区间内
|
||||
|
||||
### 4. 噪声平滑 (apply_kalman_filter)
|
||||
- **对象**:`COMMUNICATION_RANGE`(通信范围)字段
|
||||
- **方法**:使用窗口大小为 3 的移动平均进行平滑(与卡尔曼滤波效果类似的时序平滑)
|
||||
- **结果**:数值保留 2 位小数
|
||||
- **目的**:削弱测量噪声对连续数值的影响
|
||||
|
||||
### 5. 标准化与输出
|
||||
- 坐标精度统一为 2 位小数
|
||||
- 时间戳统一为 ISO-8601 格式
|
||||
- 输出清洗后 JSON 及详细报告(含各步骤的统计与示例)
|
||||
|
||||
## 快速开始
|
||||
|
||||
@@ -34,23 +59,25 @@ project/
|
||||
pip install numpy pandas
|
||||
```
|
||||
|
||||
|
||||
### 2. 运行清洗
|
||||
直接运行主程序即可:
|
||||
在项目根目录下执行:
|
||||
|
||||
```bash
|
||||
python src/main.py
|
||||
python src/cleaner.py
|
||||
```
|
||||
|
||||
程序默认读取 `data/raw_data.json`,处理后生成 `data/cleaned_data.json` 和 `report/cleaning_report.json`。
|
||||
程序会读取 `data/raw_data_sample.json`,处理后生成:
|
||||
- `data/cleaned_data_final.json`(清洗后的数据)
|
||||
- `report/detailed_cleaning_report.json`(详细清洗报告)
|
||||
|
||||
## 输出结果示例
|
||||
|
||||
**清洗前 (Raw):**
|
||||
```json
|
||||
{
|
||||
"MOBILITY": -0.5,
|
||||
"STRIKE_ACCURACY": 1.5,
|
||||
"COMMUNICATION_RANGE": 102.8116 // 含噪声
|
||||
"COMMUNICATION_RANGE": 102.8116
|
||||
}
|
||||
```
|
||||
|
||||
@@ -59,6 +86,12 @@ python src/main.py
|
||||
{
|
||||
"MOBILITY": 0.5,
|
||||
"STRIKE_ACCURACY": 1.0,
|
||||
"COMMUNICATION_RANGE": 102.81 // 平滑后
|
||||
"COMMUNICATION_RANGE": 102.81
|
||||
}
|
||||
```
|
||||
|
||||
## 清洗报告说明
|
||||
|
||||
`detailed_cleaning_report.json` 中包含:
|
||||
- **summary**:总记录数、最终记录数、去重数量
|
||||
- **details**:各字段缺失值填充数量、异常值修正数量及示例、噪声平滑处理记录、标准化说明
|
||||
|
||||
@@ -150,13 +150,18 @@ class AdvancedDataCleaner:
|
||||
print(f"完成!报告已生成至 {self.report_file}")
|
||||
|
||||
if __name__ == "__main__":
|
||||
# 使用相对路径:../data/ 表示上一级目录下的 data 文件夹
|
||||
input_path = '../data/raw_data_sample.json'
|
||||
output_path = '../data/cleaned_data_final.json'
|
||||
report_path = '../report/detailed_cleaning_report.json'
|
||||
import os
|
||||
# 基于脚本位置计算项目根目录,保证无论从哪里运行都能正确找到文件
|
||||
script_dir = os.path.dirname(os.path.abspath(__file__))
|
||||
project_root = os.path.dirname(script_dir)
|
||||
input_path = os.path.join(project_root, 'data', 'raw_data_sample.json')
|
||||
output_path = os.path.join(project_root, 'data', 'cleaned_data_final.json')
|
||||
report_path = os.path.join(project_root, 'report', 'detailed_cleaning_report.json')
|
||||
|
||||
# 确保 report 目录存在
|
||||
os.makedirs(os.path.dirname(report_path), exist_ok=True)
|
||||
|
||||
# 增加一个检查,防止路径错误
|
||||
import os
|
||||
if not os.path.exists(input_path):
|
||||
print(f"错误:找不到文件 {input_path}")
|
||||
print(f"当前工作目录是:{os.getcwd()}")
|
||||
|
||||
Reference in New Issue
Block a user