chore: PR #240 회차 2 리뷰 반영 — DOGFOOD 스니펫이 조용히 무시되는 옛 키였다

가장 나쁜 것부터.

1. 회차 1 에서 DOGFOOD §1.2 에 새로 넣은 config 스니펫이 `[image.ocr]` 였다.
   옛 키 자동 이관은 파일의 `schema_version` 이 5 보다 낮을 때만 돌고
   (`kebab-config/src/lib.rs` 의 `unwrap_or(1)` 게이트), `deny_unknown_fields`
   가 없어서 현행 v5 파일에 붙여넣으면 serde 가 통째로 버리고 경고도 안 낸다.
   `kebab doctor` 도 "config up to date" 라고 답한다. 격리 KB 로 확인: v5 +
   옛 키 → OCR 단계 자체가 없음(`parse 1ms · chunk 723ms · embed 0ms`),
   `[ingest.image.ocr]` 로 바꾸면 `ocr(ppocrv5-mobile-kor)` 가 돈다.
   그래서 같은 커밋이 추가한 1.2.f 가 no-op 이 될 뻔했다 — OCR 이 꺼지면 모든
   이미지가 "본문 0 자 + provenance 깨끗" 으로 나오는데, 이건 문서가 "글자
   없는 사진이라 정상" 이라고 읽으라고 적어 둔 바로 그 모양이다. 키를 고치고,
   같은 함정을 다음 사람이 안 밟게 §1.4 가 쓰는 형식의 주의 문단을 붙였다.
   1.2.f 에도 "진행 출력에 ocr 단계가 찍히는지 먼저 보라" 를 넣었다.

2. HOTFIXES 의 "v0.33.0 이전에 색인된" 이 한 릴리스 어긋났다. v0.33.0 태그가
   이 PR 의 base(5596d41) 자체이고 그 커밋이 `err={}` 를 담고 있다. 게다가
   스캔 PDF 페이지 렌더링(#232)이 v0.33.0 에서 처음 나갔으므로, 이 버그가
   망칠 수 있었던 스캔본은 사실상 전부 그 한 릴리스가 만든 기록이다. 문장대로
   감사하면 첫 LIKE 만 돌려 0 건을 보고 정반대 결론에 닿는다.

3. 상수 주석이 아직도 테스트가 위쪽 방향까지 잡는다고 말했다. 회차 1 은
   문장을 길게 바꿨을 뿐 주장은 그대로 뒀다. 위쪽은 `const _` 가 잡고 테스트는
   못 잡는다는 걸 그대로 적었다.

4. `{e:#}` 를 고정하는 테스트가 없었다. mock 오류가 단층이라 anyhow 가 `{}` 와
   `{:#}` 에서 똑같이 찍었고, 그래서 되돌려도 통과했다. mock 을 실제와 같은 두
   층으로 만들고 안쪽 원인을 단언한다. `{}` 로 되돌리면 실패하는 것 확인.

나머지:
- `pdf_ocr_apply.rs:475` 가 회차 1 반영 커밋 자신에 밀려 어긋났다 (같은 종류가
  한 PR 에서 두 번). 줄 번호를 빼고 함수명으로 고정했다.
- const-assert 주석이 안 잰 구간(17~39)을 잰 것처럼 말했다. 게다가 17 로
  올리면 새로 버려지는 건 폭 16 — 직접 비어 있다고 잰 폭이다.
- "재처리 대상은 PDF 전부" 정정이 HOTFIXES 에서 멈춰 ARCHITECTURE 와 컴포넌트
  README 에는 아직 "스캔본" 이라고 적혀 있었다.
- 컴포넌트 README 의 `engine_id() / run(...)` 은 존재한 적 없는 시그니처다.
  산문과 mermaid 다이어그램 양쪽을 실제 trait 에 맞췄다.
- §1.2 verify 의 `err=rec session run` 은 이미지 경로가 안 내는 문자열이다
  (KB 전체 집계용을 이미지 전용 절에 복사해 왔다).
- 영향 문서 집계 쿼리에 실행 시점이 빠져 있었다. bump 가 documents 행을
  purge 하므로 한 번 색인한 뒤에는 항상 0 이 나온다. "색인 전에 세라" 와,
  이미 색인했을 때 쓸 `pdf_ocr_events` 대안을 적었다.
- Gitea PR 본문을 현재 브랜치에 맞게 다시 썼다.

알고 남긴 것: 커밋 d6654ab 메시지의 "실제 하한은 4 였다" 는 틀렸다 (하한은 5,
4 는 가장 큰 실패 폭 — 같은 메시지 두 문단 뒤와 자기모순). 코드와 문서는 전부
정확하다. 고치려면 리뷰 중인 브랜치에 강제 푸시가 필요해서 두었다.

검증: 워크스페이스 1301 passed / 0 failed, clippy -D warnings 무경고.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-28 18:08:38 +09:00
parent 77ed432af8
commit b60148b123
7 changed files with 51 additions and 24 deletions

View File

@@ -1,6 +1,6 @@
use std::sync::Mutex;
use anyhow::Result;
use anyhow::{Context, Result};
use kebab_core::{Lang, OcrText};
use kebab_parse_image::OcrEngine;
@@ -52,7 +52,13 @@ impl OcrEngine for MockOcrEngine {
fn recognize(&self, _img: &[u8], _hint: Option<&Lang>) -> Result<OcrText> {
if self.fail {
anyhow::bail!("mock failure");
// Layered on purpose: the real paddle-onnx failure arrives as an
// ORT message under a `.context("rec session run")`, and anyhow's
// plain Display shows only the outer layer. A single-layer error
// would render identically under `{}` and `{:#}`, so it could not
// tell whether the provenance note keeps the cause (issue #239).
return Err(anyhow::anyhow!("mock inner cause"))
.context("mock failure");
}
let mut idx = self.call_index.lock().unwrap();
let text = self

View File

@@ -324,6 +324,18 @@ fn ocr_engine_failure_surfaces_as_warning() {
warning_with_failure,
"OCR failure 의 error message 가 warning event 의 note 안"
);
// issue #239: the note must carry the whole error chain, not just the
// outermost layer. The real cause (ORT's "Invalid input shape") sits under
// a `.context`, so a note formatted with `{e}` instead of `{e:#}` drops it
// and the KB can no longer be searched for which documents were hit.
let warning_with_cause = canonical.provenance.events.iter().any(|e| {
e.kind == kebab_core::ProvenanceKind::Warning
&& e.note.as_deref().unwrap_or("").contains("mock inner cause")
});
assert!(
warning_with_cause,
"provenance note 가 error chain 의 안쪽 원인까지 담아야 한다 (`{{e:#}}`)"
);
}
// Test 9: dual-block ordinals are deterministic and unique

View File

@@ -57,16 +57,18 @@ const REC_HEIGHT: u32 = 48;
/// `Invalid input shape: {1,0}` — taking every already-recognized box on the
/// image down with it (issue #239). Measured against the bundled
/// `korean_ppocrv5_mobile_rec.onnx`: 1..=4 always fail, 5.. always succeed.
/// `rec_min_width_is_the_graph_floor` holds the constant from both
/// directions: too low and the graph rejects it, too high and it starts
/// discarding crops the graph would have accepted.
/// `rec_min_width_is_the_graph_floor` holds the constant from below — set it
/// under the graph's floor and that test errors. The upward direction cannot
/// be a test (both of its assertions pass *better* as the constant grows), so
/// it is the `const _` ceiling right below instead.
const REC_MIN_WIDTH: u32 = 5;
/// Raising `REC_MIN_WIDTH` is only free while it stays inside the band that
/// was measured to decode nothing anyway: widths 5..=16 come back empty even
/// with real ink in them, while width 40 reads glyphs at 0.97+ confidence
/// (issue #239). Above 16 the guard starts discarding crops the graph would
/// have read — the same silent loss this constant exists to prevent — so the
/// ceiling is enforced at compile time rather than left to review.
/// Raising `REC_MIN_WIDTH` is only known to be free while it stays inside the
/// band that was actually measured: widths 5..=16 come back empty even with
/// real ink in them, while width 40 reads glyphs at 0.97+ confidence (issue
/// #239). 17..=39 was never swept, so a ceiling above 16 is not measurement
/// any more — and by 40 the guard would be discarding crops the graph reads,
/// which is the silent loss this constant exists to prevent. Enforced at
/// compile time rather than left to review.
const _: () = assert!(
REC_MIN_WIDTH <= 16,
"REC_MIN_WIDTH is past the measured no-loss ceiling (16)"

View File

@@ -24,7 +24,7 @@ Cargo workspace, 함수 호출 기반 모듈러 모놀리스. UI binary (`kebab-
| OCR (PDF, v0.20.0+) | Ollama vision LM (default `qwen2.5vl:3b`) — post-extract enrichment via `kebab-app::pdf_ocr_apply` (H-1 resolution). DCTDecode-only v1 (FlateDecode/CCITTFax skip + warning). family asymmetry vs image OCR: PoC alnum 94.79% (qwen2.5vl) >> 27% (gemma4:e4b 받침), 본 단계에서 PDF OCR 만 qwen2.5vl. |
| Image caption | Ollama vision LM, runtime gate `image.caption.enabled` (default OFF) |
| RAG groundedness 검증 | `kebab-nli` 의 mDeBERTa-v3 XNLI 가 `(packed_chunks, generated_answer)` entailment 검사 (fb-41). `[rag] nli_threshold > 0` (default 0 = disabled, production 권장 0.5) 일 때 활성 — 미달 시 `refusal_reason = nli_verification_failed` (LLM self-judge ceiling 보완). 첫 호출 시 ~280 MB ONNX 자동 다운로드 |
| PDF parser | `lopdf` per-page 텍스트 + 스캔 페이지 래스터화. 래스터는 **pdfium 페이지 렌더링**(`page_render::PageRenderer`, issue #232) 이 1순위 — 필터·XObject 구성과 무관하게 페이지를 그린다. pdfium 이 없으면 `page_image::extract_dctdecode_page_image` 로 떨어지며 그 경우 단일 DCTDecode 이미지 페이지만 OCR 된다. pdfium 은 공유 라이브러리로만 배포돼 링크하면 단일 바이너리 원칙이 깨지므로 **런타임 바인딩**이고, `[ingest.pdf.ocr] render_library` 로 경로를 지정하거나 로더 경로에 두면 된다. `kebab doctor` 의 `pdf_render` 가 어느 쪽인지 보고한다. `chunker_version = "pdf-page-v1"` 하드코딩 (HOTFIXES P7-3). `parser_version = "pdf-text-v3"` (issue #232 에서 v1 → v2, issue #239 에서 v2 → v3 — 둘 다 기존 색인 스캔본 재처리 유발). |
| PDF parser | `lopdf` per-page 텍스트 + 스캔 페이지 래스터화. 래스터는 **pdfium 페이지 렌더링**(`page_render::PageRenderer`, issue #232) 이 1순위 — 필터·XObject 구성과 무관하게 페이지를 그린다. pdfium 이 없으면 `page_image::extract_dctdecode_page_image` 로 떨어지며 그 경우 단일 DCTDecode 이미지 페이지만 OCR 된다. pdfium 은 공유 라이브러리로만 배포돼 링크하면 단일 바이너리 원칙이 깨지므로 **런타임 바인딩**이고, `[ingest.pdf.ocr] render_library` 로 경로를 지정하거나 로더 경로에 두면 된다. `kebab doctor` 의 `pdf_render` 가 어느 쪽인지 보고한다. `chunker_version = "pdf-page-v1"` 하드코딩 (HOTFIXES P7-3). `parser_version = "pdf-text-v3"` (issue #232 에서 v1 → v2, issue #239 에서 v2 → v3). 고친 것은 스캔본 경로지만 **재처리 대상은 기존 색인 PDF 전부** 다 — base `parser_version` 이 `id_for_doc` 에 접히므로 doc_id 가 바뀌고 store 가 다시 쓰인다 (HOTFIXES 2026-08-28). |
| code parser | `tree-sitter` + `tree-sitter-rust` / `tree-sitter-python` / `tree-sitter-typescript` / `tree-sitter-javascript` / `tree-sitter-go` / `tree-sitter-java` / `tree-sitter-kotlin-ng` — **parser-side** (`kebab-parse-code`), chunker-side 아님 (design §6.3). chunker versions: Rust = `code-rust-ast-v1`, Python = `code-python-ast-v1`, TypeScript = `code-ts-ast-v1`, JavaScript = `code-js-ast-v1`, Go = `code-go-ast-v1`, Java = `code-java-ast-v1`, Kotlin = `code-kotlin-ast-v1`. (v0.32.0 #220: 9개 언어 chunker 가 단일 `CodeAstV1Chunker` 로 통합 — `for_lang(lang)` 가 per-lang `chunker_version` 라벨을 verbatim 매핑. chunker 는 tree-sitter 미사용·`lang` 은 `SourceSpan::Code` 데이터에서 흐르므로 9개 struct 차이는 `VERSION_LABEL` 문자열뿐이었음 → chunk_id byte-identical.) `ast_chunk_max_lines = 200` 상수 고정 (HOTFIXES 2026-05-19 — Chunker trait 이 per-medium config 미노출). Kotlin grammar 은 `tree-sitter-kotlin-ng` 사용 — bare `tree-sitter-kotlin` 은 tree-sitter 0.21–0.23 에 고착되어 있어 사용 불가. **Tier 2 (p10-2)**: YAML/k8s → `serde_yaml_ng` + `k8s-manifest-resource-v1` (apiVersion+kind per resource), Dockerfile → `dockerfile-file-v1` (whole-file), Cargo.toml/go.mod/.json/.xml/.groovy → `manifest-file-v1` (whole-file). Tier 2 chunkers live in `kebab-chunk`; no tree-sitter grammar needed (structure from file type, not AST). **Tier 3 (p10-3)**: shell scripts (`.sh`/`.bash`/`.zsh`) direct → `code-text-paragraph-v1` (blank-line paragraph segmentation + 80-line / 20-overlap line-window for oversize). Same chunker also serves as fallback when Tier 1/2 emit 0 chunks or Err — non-k8s YAML / invalid YAML / AST extractor failures all picked up. symbol = None; lang preserved from input doc. **Tier 1 family complete (p10-1D)**: C (`tree-sitter-c`, `code-c-ast-v1`, `.c`/`.h`) + C++ (`tree-sitter-cpp`, `code-cpp-ast-v1`, `.cpp`/`.cc`/`.cxx`/`.hpp`/`.hh`/`.hxx`). C symbol = function name only; C++ symbol = `namespace::Class::method` (recursive nesting). `.h` 가 C++ syntax 만나면 tree-sitter-c parse 실패 → Tier 3 fallback. |
| symbol path 형식 | workspace path → module path: Python = dotted prefix (`kebab_eval.metrics.compute_mrr`), TypeScript/JavaScript = slash-style prefix (`src/Foo.Foo.search`), Go = `package.Func` / `package.(*Receiver).Method`, Java/Kotlin = `com.foo.Foo.bar` (패키지+클래스+메서드/필드), C = 함수명, C++ = `namespace::Class::method`. Rust 1A-2 는 file-scope nesting 만 (workspace prefix 없음, 비일관 수용 — HOTFIXES 2026-05-20). code chunk 은 `citation.kind = "code"` + `citation.lang` + `symbol` + line range, SearchHit 에 `code_lang` + `repo`(`.git` walk-up 디렉토리명) backfill. |
| Desktop | Tauri 2 + `pdfjs-dist` (native PDF render backend 금지) — P9-5 |

View File

@@ -122,9 +122,9 @@ endpoint = "http://192.168.0.47:11434"
enabled = false # opt-in
```
`paddle-onnx` 백엔드 (v0.27.0~, in-process ONNX — Ollama 없이 돈다):
**Config (v0.28.0~)**: 위 블록은 `[ingest.image.ocr]` / `[ingest.image.caption]` 로 옮겨졌다. 옛 키를 그대로 쓰면 **`schema_version` 이 5 보다 낮은 파일에서만** 로드 시 자동 이관된다 — 이미 `schema_version = 5` 인 config 에 `[image.ocr]` 를 붙여 넣으면 경고 없이 통째로 무시되고 OCR 이 꺼진 채로 돈다. `paddle-onnx` 백엔드 (v0.27.0~, in-process ONNX — Ollama 없이 돈다):
```toml
[image.ocr]
[ingest.image.ocr]
enabled = true
engine = "paddle-onnx"
```
@@ -132,8 +132,8 @@ engine = "paddle-onnx"
**verify**:
- `*.png` / `*.jpg` / `*.jpeg` 만 ingest target.
- OCR text 가 `Block::ImageRef.ocr.joined` 안.
- `[image.caption].enabled=true` 시 caption 도.
- (issue #239) 한 장도 **통째로** 비지 않는다. 얇은 검출 박스 하나가 rec 세션을 실패시키면 그 이미지의 인식 결과가 전량 버려지던 버그였다. 색인은 성공으로 끝나므로 `documents.provenance_json` 에 `Invalid input shape` 또는 `err=rec session run` 이 남았는지로 확인한다.
- caption 토글을 켜면 caption 도.
- (issue #239) 한 장도 **통째로** 비지 않는다. 얇은 검출 박스 하나가 rec 세션을 실패시키면 그 이미지의 인식 결과가 전량 버려지던 버그였다. 색인은 성공으로 끝나므로 `documents.provenance_json` 에 `Invalid input shape` 이 남았는지로 확인한다. (PDF 경로는 노트 형식이 달라 문자열이 다르다 — HOTFIXES 2026-08-28 참고.)
**scenarios**:
- 1.2.a Korean OCR (한국어 scan PNG) → OCR text + search hit.
@@ -141,7 +141,7 @@ engine = "paddle-onnx"
- 1.2.c photo (자연 사진, OCR 없음) → empty OCR or warning.
- 1.2.d corrupt image → graceful error.
- 1.2.e oversized image (> max_pixels) → downscale.
- 1.2.f (issue #239) `engine = "paddle-onnx"` 로 §13.4 이미지 코퍼스 전량 → OCR 오류 0 건. 본문이 파일명뿐인 문서가 남으면 provenance 를 확인한다. 글자가 없는 사진이라 0 자인 것과 이 버그로 통째로 버려진 것은 다르다 — 후자만 provenance 에 오류가 남는다.
- 1.2.f (issue #239) `[ingest.image.ocr] engine = "paddle-onnx"` 로 §13.4 이미지 코퍼스 전량 → OCR 오류 0 건. 먼저 진행 출력에 `ocr(ppocrv5-mobile-kor…)` 단계가 실제로 찍히는지 본다 — 안 찍히면 OCR 이 꺼진 것이고, 그 상태에서는 모든 이미지가 본문 0 자로 나와 시나리오가 통과한 것처럼 보인다. 본문이 파일명뿐인 문서가 남으면 provenance 를 확인한다. 글자가 없는 사진이라 0 자인 것과 이 버그로 통째로 버려진 것은 다르다 — 후자만 provenance 에 오류가 남는다.
### §1.3 PDF text ingest (P7-1)

View File

@@ -36,8 +36,9 @@ classDiagram
}
class OcrEngine {
<<trait kebab-parse-image>>
engine_id() str
run(image_bytes, langs) OcrText
engine_name() str
engine_version() String
recognize(image_bytes, lang_hint) Result~OcrText~
}
class OllamaVisionOcr {
endpoint, model, max_pixels
@@ -106,13 +107,13 @@ flowchart LR
**PDF** (`kebab-parse-pdf`):
- `PdfTextExtractor` — `Extractor` 구현체. `lopdf::Document::load_mem` 로 한 번 파싱, encrypted 면 즉시 bail.
- `PARSER_VERSION = "pdf-text-v3"` — version cascade entry (issue #232 에서 v1 → v2, 페이지 렌더링 도입으로 기존 색인 스캔본 재처리 유발; issue #239 에서 v2 → v3, 얇은 검출 박스가 페이지 OCR 을 통째로 날리던 것을 고치면서 기존 색인 스캔본 재처리 유발). (HOTFIXES P7-2 의 chunker_version `pdf-page-v1` 와 별개.)
- `PARSER_VERSION = "pdf-text-v3"` — version cascade entry (issue #232 에서 v1 → v2, 페이지 렌더링 도입; issue #239 에서 v2 → v3, 얇은 검출 박스가 페이지 OCR 을 통째로 날리던 것을 고치면서). 고친 것은 둘 다 스캔본 경로지만 **재처리 대상은 기존 색인 PDF 전부** — base 가 `id_for_doc` 에 접혀 doc_id 가 바뀐다. (HOTFIXES P7-2 의 chunker_version `pdf-page-v1` 와 별개.)
- 빈 페이지 / extract 실패 → `Block::Paragraph` 빈 inlines + `ProvenanceKind::Warning("scanned candidate")`. OCR fallback 미구현.
**Image** (`kebab-parse-image`):
- `ImageExtractor` — `Extractor` 구현체. `MAX_DECODE_DIM = 16384` 초과 거부 (decode bomb 방어).
- `PARSER_VERSION = "image-meta-v2"` — version cascade entry (issue #239 에서 v1 → v2, 얇은 검출 박스가 이미지 OCR 을 통째로 날리던 것을 고치면서 기존 색인 이미지 재처리 유발).
- `OcrEngine` (trait) — `engine_id() / run(...) -> OcrText`. `OcrText.engine` 필드로 trust level 분기.
- `PARSER_VERSION = "image-meta-v2"` — version cascade entry (issue #239 에서 v1 → v2, 얇은 검출 박스가 이미지 OCR 을 통째로 날리던 것을 고치면서). PDF 와 마찬가지로 기존 색인 이미지 **전부** 가 재처리 대상이다.
- `OcrEngine` (trait) — `engine_name() -> &'static str` / `engine_version() -> String` / `model() -> &str` / `recognize(&[u8], Option<&Lang>) -> Result<OcrText>`. `OcrText.engine` 필드로 trust level 분기.
- `OllamaVisionOcr { endpoint, model, max_pixels }` — `ollama-vision` 백엔드 (기본값). `apply_ocr(block, engine, langs)` 가 `ImageRefBlock.ocr` 슬롯 채움.
- `OnnxPaddleOcr` — `paddle-onnx` 백엔드 (v0.27.0, PP-OCRv5 ONNX in-process). rec 세션 입력 폭 하한은 `REC_MIN_WIDTH = 5` — 그 아래는 세션에 넣지 않고 빈 문자열을 돌려준다 (issue #239).
- `caption_image(lm: &dyn LanguageModel, prep, opts) -> Result<ModelCaption>` — `LanguageModel.generate_stream` 의 vision 입력 (`GenerateRequest.images`) 사용. `apply_caption` 이 block 에 in-place 주입.

View File

@@ -54,6 +54,8 @@ git history.
모델 에셋이 in-tree 로 커밋돼 있으므로(`git ls-files crates/kebab-parse-image/assets/`) skip 가드는 붙이지 않았다 — 조용히 안 도는 테스트가 이 이슈가 경고하는 바로 그 함정이다.
PDF 노트의 `{e:#}` 도 `ocr_engine_failure_surfaces_as_warning` 이 고정한다. 원래 이 테스트는 mock 이 단층 오류를 내서 `{}` 로 되돌려도 통과했다 — anyhow 는 원인이 없는 오류를 두 형식에서 똑같이 찍기 때문이다. mock 을 실제와 같은 두 층 오류로 바꾸고 안쪽 원인까지 단언하도록 했다. `{}` 로 되돌리면 실패한다.
### 실측 (도그푸딩 말뭉치 이미지 240 개 파일)
`corpus/images/` 전량(jpg 172 · png 63 · jpeg 4 · tif 1 = 240 개 파일. `kebab ingest` 가 집는 것은 tif 를 뺀 239 개지만, 여기서는 OCR 엔진을 파일 목록에 직접 물렸다)을 수정 전후로 같은 엔진(`ppocrv5-mobile-kor-1b55f062d055`)·같은 설정(score 0.3 / unclip 1.5 / max_boxes 1000 / max_pixels 2048)으로 돌렸다.
@@ -75,7 +77,11 @@ git history.
부작용이 없다는 것도 확인했다: 원래 성공하던 204 장의 인식 글자 수가 **한 장도 변하지 않았다**. 이 수정은 순수 가산이다.
36 장 중 4 장은 수정 후에도 0 자인데, 글자가 없는 사진(예: `photos/Aphid_2007_1.jpg`, 진딧물 접사)이라 정상이다. 이 이슈로 인한 손실과 "원래 글자가 없어서 비는 문서"를 혼동하면 안 된다 — 성공한 204 장 중에서도 31 장은 det 가 박스를 못 찾아 정상적으로 0 자다. KB 쪽에서 영향 문서를 셀 때는 본문 길이가 아니라 `provenance_json LIKE '%Invalid input shape%'` 로 걸러야 한다. 단, **v0.33.0 이전에 색인된 스캔 PDF 는 이 필터에 안 걸린다** — PDF 경로가 노트를 `err={}` 로 찍었고 anyhow 의 기본 Display 는 가장 바깥 context 하나(`rec session run`)만 내보내기 때문이다. 이미지 경로는 처음부터 `{err:#}` 라 체인 전체가 들어간다. 이번에 PDF 쪽도 `{e:#}` 로 맞춰 두 경로의 노트 형식을 통일했으므로 앞으로 색인되는 것은 양쪽 다 걸린다. 이전 색인분까지 세려면 `OR provenance_json LIKE '%err=rec session run%'` 을 함께 걸어야 한다.
36 장 중 4 장은 수정 후에도 0 자인데, 글자가 없는 사진(예: `photos/Aphid_2007_1.jpg`, 진딧물 접사)이라 정상이다. 이 이슈로 인한 손실과 "원래 글자가 없어서 비는 문서"를 혼동하면 안 된다 — 성공한 204 장 중에서도 31 장은 det 가 박스를 못 찾아 정상적으로 0 자다. KB 쪽에서 영향 문서를 셀 때는 본문 길이가 아니라 `provenance_json LIKE '%Invalid input shape%'` 로 걸러야 한다. 다만 이 쿼리에는 조건이 둘 붙는다.
**언제 세느냐.** 새 바이너리로 `kebab ingest` 를 돌리기 **전에** 세야 한다. 아래 §재색인 의 `parser_version` bump 가 해당 경로의 documents 행을 지우고 다시 쓰므로, 한 번 색인한 뒤에는 이 쿼리가 0 을 돌려준다. "업그레이드 → 색인 → 릴리스 노트 읽기" 순서로 가면 안 당한 것처럼 보인다. 이미 색인해 버렸다면 PDF 쪽은 `SELECT count(*) FROM pdf_ocr_events WHERE success = 0 AND reason = 'ocr_error'` 로 아직 셀 수 있다 — `pdf_ocr_events` 는 documents 에 FK 가 없어 purge 를 넘겨 살아남는다(`logging.retention_days` 기본 30 일 prune 만 받는다).
**스캔 PDF 는 이 필터로 안 걸린다.** 이번 수정 이전에 나간 **모든** 릴리스에서 — v0.33.0 을 포함해서 — PDF 경로는 노트를 `err={}` 로 찍었고, anyhow 의 기본 Display 는 가장 바깥 context 하나(`rec session run`)만 내보낸다. 하필 스캔 PDF 페이지 렌더링(#232) 자체가 v0.33.0 에서 처음 나갔으므로, 이 버그가 망칠 수 있었던 스캔본은 사실상 전부 그 한 릴리스가 만든 기록이다. 그래서 스캔본까지 세려면 `OR provenance_json LIKE '%err=rec session run%'` 을 반드시 함께 걸어야 한다. 이미지 경로는 처음부터 `{err:#}` 라 체인 전체가 들어갔고, 이번에 PDF 쪽도 `{e:#}` 로 맞췄으므로 앞으로 색인되는 것은 양쪽 다 첫 필터에 걸린다.
### 재색인: 버전 두 개를 올렸다
@@ -84,7 +90,7 @@ git history.
- `image-meta-v1` → **`image-meta-v2`**
- `pdf-text-v2` → **`pdf-text-v3`** (스캔 PDF 도 같은 `run_rec` 을 타므로 같은 손실을 겪었다)
**비싼 OCR 은 대부분 다시 안 돈다.** OCR 산출물은 `derivation_cache` 에 **소스 바이트** 키로 들어가 있어서 `parser_version` 캐스케이드와 분리돼 있다(`docs/ARCHITECTURE.md:35` 의 derivation_cache 행, v0.31.0 #217). 그리고 실패한 OCR 은 캐시에 **저장되지 않는다** — `Err` 분기가 `derivation_cache_put` 앞에서 빠져나간다(이미지 `ingest.rs:1750`, PDF `pdf_ocr_apply.rs:475`). 그래서 이미 성공했던 문서는 캐시에 히트해 엔진 호출을 건너뛰고, **실제로 다시 OCR 되는 건 이 버그로 실패했던 문서뿐**이다.
**비싼 OCR 은 대부분 다시 안 돈다.** OCR 산출물은 `derivation_cache` 에 **소스 바이트** 키로 들어가 있어서 `parser_version` 캐스케이드와 분리돼 있다(`docs/ARCHITECTURE.md:35` 의 derivation_cache 행, v0.31.0 #217). 그리고 실패한 OCR 은 캐시에 **저장되지 않는다** — `Err` 분기가 `derivation_cache_put` 앞에서 빠져나간다(이미지는 `ingest.rs` 의 `ingest_one_image_asset`, PDF 는 `pdf_ocr_apply.rs` 의 `apply_ocr_to_pdf_pages`). 줄 번호를 안 적은 건 이 항목이 처음 썼던 두 참조가 같은 PR 의 후속 커밋에 밀려 둘 다 어긋났기 때문이다. 그래서 이미 성공했던 문서는 캐시에 히트해 엔진 호출을 건너뛰고, **실제로 다시 OCR 되는 건 이 버그로 실패했던 문서뿐**이다.
**하지만 나머지 비용은 전부 든다.** `id_for_doc` 이 접는 것은 composite 가 아니라 **base** PARSER_VERSION 이므로(`kebab-parse-image/src/lib.rs`, `kebab-parse-pdf/src/lib.rs` 의 `extract`; composite 는 그 뒤에 `canonical.parser_version` 에만 찍힌다) 이 bump 는 **모든 이미지·PDF 문서의 doc_id 를 바꾼다**. 그러면 `ingest.rs` 의 `purge_workspace_path_for_parser_bump` 가 돌아 documents 행이 지워지고(blocks·chunks·embedding_records CASCADE) 해당 chunk_id 의 Lance 벡터도 전량 삭제된 뒤, 재파싱·재청킹·재임베딩·재삽입이 이어진다. 임베딩은 파생물 캐시에 히트하지만 행은 다시 쓴다. doc_id 는 wire 필수 필드이자 `kebab inspect doc <id>` 의 핸들이라, 파일을 하나도 안 고쳤는데 전부 한꺼번에 바뀐다.