기준 데이터 정렬 후 빼기
설명
측정 데이터와 기준(Reference) 파일을 유효 구간·중앙값으로 맞춘 뒤 차이를 구한다
필요 데이터
현재 불러온 측정 데이터 + 같은 컬럼 이름을 가진 기준 파일 1개
출력
컬럼별 (측정 − 기준) 차이 표 + 비교 그래프
필요 패키지
- numpy
- pandas
- matplotlib
태그
science, visualization
작성자
SDSL | MDS 1.0
측정 데이터와 기준(Reference) 파일을 유효 구간·중앙값으로 맞춘 뒤 차이를 구한다
Aligns measurement and reference data by valid range and median, then subtracts
README
측정 데이터와 기준(Reference) 파일을 유효 구간·중앙값으로 맞춘 뒤 차이를 구한다
현재 불러온 측정 데이터 + 같은 컬럼 이름을 가진 기준 파일 1개
컬럼별 (측정 − 기준) 차이 표 + 비교 그래프
science, visualization
SDSL | MDS 1.0
# MDS_TITLE: 기준 데이터 정렬 후 빼기
# MDS_REQUIRES: numpy, pandas, matplotlib
# MDS_DESC: 측정 데이터와 기준(Reference) 파일을 유효 구간·중앙값으로 맞춘 뒤 차이를 구한다
# MDS_DESC[en]: Aligns measurement and reference data by valid range and median, then subtracts
# MDS_TAGS: science, visualization
# MDS_INPUT: 현재 불러온 측정 데이터 + 같은 컬럼 이름을 가진 기준 파일 1개
# MDS_OUTPUT: 컬럼별 (측정 − 기준) 차이 표 + 비교 그래프
# MDS_AUTHOR: SDSL
# MDS_VERSION: 1.0
def valid_range(values: np.ndarray, threshold: float) -> tuple[int, int] | None:
idx = np.where(values > threshold)[0]
if idx.size == 0:
return None
return int(idx[0]), int(idx[-1])
def aligned_difference(meas: pd.Series, ref: pd.Series, threshold: float) -> np.ndarray:
v1 = pd.to_numeric(meas, errors='coerce').dropna().to_numpy()
v2 = pd.to_numeric(ref, errors='coerce').dropna().to_numpy()
r1, r2 = valid_range(v1, threshold), valid_range(v2, threshold)
if r1 is None or r2 is None:
return np.array([])
n = min(r1[1] - r1[0] + 1, r2[1] - r2[0] + 1)
a = v1[r1[0]:r1[0] + n]
b = v2[r2[0]:r2[0] + n]
diff = a - (b - (np.median(b) - np.median(a)))
return diff[2:-2] if diff.size > 4 else diff
def subtract_reference(df: pd.DataFrame, ref_df: pd.DataFrame, threshold: float) -> pd.DataFrame:
dff = df.copy()
common = [c for c in dff.columns if c in ref_df.columns]
parts = {c: pd.Series(aligned_difference(dff[c], ref_df[c], threshold)) for c in common}
return pd.DataFrame(parts)
def plot_difference(table: pd.DataFrame) -> None:
try:
plt.style.use('seaborn-v0_8-whitegrid')
except OSError:
plt.style.use('seaborn-whitegrid')
plt.rcParams['font.family'] = 'Malgun Gothic'
plt.rcParams['axes.unicode_minus'] = False
fig, ax = plt.subplots(figsize=(11, 5))
for col in table.columns[:8]:
ax.plot(table.index, table[col], linewidth=1.2, label=str(col))
ax.axhline(0, color='#888888', linewidth=0.8)
ax.set_title('측정 − 기준 차이', fontsize=12, fontweight='semibold')
ax.set_xlabel('정렬된 샘플 순번')
ax.set_ylabel('차이')
ax.legend(fontsize=8)
ax.grid(True, alpha=0.3)
plt.tight_layout()
plt.show()
return None
def run_reference_subtract(df: pd.DataFrame) -> pd.DataFrame | None:
ref_path = mIO.open_file_single(dialog_title='기준(Reference) 파일 선택',
file_filter='데이터 파일 (*.csv *.xlsx *.xls);;모든 파일 (*.*)')
if ref_path is None:
return None
ref_df = mIO.read_from_file(ref_path)
if ref_df is None:
mIO.print('기준 파일을 읽지 못했습니다.')
return None
table = subtract_reference(df, ref_df, threshold=0.001)
if table.empty:
mIO.print('측정 데이터와 기준 파일에 공통 컬럼이 없습니다.')
return None
plot_difference(table)
return table
result = run_reference_subtract(df)