회차 3 은 3 건으로 수렴했고 그중 둘이 회차 2 가 만든 것이다. 1. 릴리스 경계를 바로잡다가 반대쪽으로 과하게 밀었다. "이 버그가 망칠 수 있었던 스캔본은 사실상 전부 v0.33.0 이 만든 기록" 이라고 썼는데, 페이지 렌더링(#232)이 v0.33.0 신규인 건 맞지만 스캔 PDF OCR 자체는 그보다 오래됐다. `build_pdf_ocr_engine` 의 paddle 분기를 태그별로 세면 v0.28.0=0, v0.30.0=0, v0.31.0=1, v0.32.0=1 이다. v0.31.0 부터 PDF OCR 엔진으로 paddle-onnx 를 고를 수 있었고 그때 `run_rec` 에는 폭 가드가 없었다. 다만 v0.32.0 까지는 단일 DCTDecode 이미지 페이지만 대상이었다. #232 는 래스터 경로를 임의 인코딩까지 넓힌 것이지 만든 것이 아니다. (리뷰는 v0.28.0 이라고 했으나 태그별로 직접 세어 v0.31.0 로 적었다.) 2. 색인 후에 쓸 대안 쿼리의 단위가 틀렸다. 앞 문장은 영향 **문서** 수를 세라는데 `pdf_ocr_events` 는 페이지마다 한 행이고 유니크 제약도 없다 — 40 쪽을 잃은 문서 하나가 40 을 더하고, 수정 전 색인을 두 번 돌렸으면 또 두 배다. `count(DISTINCT doc_id)` 가 둘 다 접는다(이 경로는 doc_id 를 무조건 넘기므로 NULL 로 빠지는 행이 없다). 그리고 `'ocr_error'` 는 `recognize()` 의 모든 실패를 받는 통칭이라 기본 엔진 ollama-vision 을 쓰는 KB 도 #239 와 무관한 행을 쌓는다. V008 에 `ocr_engine` 컬럼이 이미 있으니 조건으로 거른다. paddle-onnx 안에서도 상한이라는 점은 문장으로 밝혔다. 3. ARCHITECTURE.md 에서 이 PR 이 다시 쓴 줄이 하필 이 PR 이 "조용히 무시된다" 고 문서화한 옛 키(`[image.ocr] engine`)를 쓰고 있었다. 같은 표 바로 아래 PDF 행은 이미 `[ingest.pdf.ocr]` 형태다. 곁다리: 회차 3 델타가 mock_ocr.rs 에 rustfmt 어긋남을 1 건 만들었다 (main 0 → 1). 이 PR 은 새 어긋남 0 을 유지해 왔으므로 되돌렸다. 검증: 워크스페이스 1301 passed / 0 failed, clippy -D warnings 무경고, 손댄 파일의 rustfmt 어긋남이 main 과 동일(새로 만든 것 0). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
77 lines
2.2 KiB
Rust
77 lines
2.2 KiB
Rust
use std::sync::Mutex;
|
|
|
|
use anyhow::{Context, Result};
|
|
use kebab_core::{Lang, OcrText};
|
|
use kebab_parse_image::OcrEngine;
|
|
|
|
pub struct MockOcrEngine {
|
|
expected_texts: Vec<String>,
|
|
call_index: Mutex<usize>,
|
|
fail: bool,
|
|
}
|
|
|
|
impl MockOcrEngine {
|
|
/// Single text (backward-compat ctor for pdf_ocr_apply.rs 10 sites).
|
|
pub fn single(text: impl Into<String>, fail: bool) -> Self {
|
|
Self {
|
|
expected_texts: vec![text.into()],
|
|
call_index: Mutex::new(0),
|
|
fail,
|
|
}
|
|
}
|
|
|
|
/// Per-page texts (cursor advances per recognize call).
|
|
pub fn per_page(texts: Vec<String>, fail: bool) -> Self {
|
|
Self {
|
|
expected_texts: texts,
|
|
call_index: Mutex::new(0),
|
|
fail,
|
|
}
|
|
}
|
|
|
|
/// Number of `recognize` calls so far (cache-hit tests assert this stays
|
|
/// flat across a re-run).
|
|
pub fn call_count(&self) -> usize {
|
|
*self.call_index.lock().unwrap()
|
|
}
|
|
}
|
|
|
|
impl OcrEngine for MockOcrEngine {
|
|
fn engine_name(&self) -> &'static str {
|
|
"mock-ocr"
|
|
}
|
|
|
|
fn engine_version(&self) -> String {
|
|
"mock-v1".to_string()
|
|
}
|
|
|
|
#[allow(clippy::unnecessary_literal_bound)]
|
|
fn model(&self) -> &str {
|
|
"mock-model"
|
|
}
|
|
|
|
fn recognize(&self, _img: &[u8], _hint: Option<&Lang>) -> Result<OcrText> {
|
|
if self.fail {
|
|
// Layered on purpose: the real paddle-onnx failure arrives as an
|
|
// ORT message under a `.context("rec session run")`, and anyhow's
|
|
// plain Display shows only the outer layer. A single-layer error
|
|
// would render identically under `{}` and `{:#}`, so it could not
|
|
// tell whether the provenance note keeps the cause (issue #239).
|
|
return Err(anyhow::anyhow!("mock inner cause")).context("mock failure");
|
|
}
|
|
let mut idx = self.call_index.lock().unwrap();
|
|
let text = self
|
|
.expected_texts
|
|
.get(*idx)
|
|
.cloned()
|
|
.unwrap_or_else(|| self.expected_texts.last().cloned().unwrap_or_default());
|
|
*idx += 1;
|
|
Ok(OcrText {
|
|
joined: text,
|
|
regions: vec![],
|
|
engine: "mock-ocr".to_string(),
|
|
engine_version: "mock-v1".to_string(),
|
|
})
|
|
}
|
|
}
|