feat(kebab-parse-image): P6-2 OCR adapter — Ollama-vision default
- 새 모듈 `crates/kebab-parse-image/src/ocr.rs` 추가. spec 의 `OcrEngine`
trait 그대로 + `OllamaVisionOcr` default 구현 + `apply_ocr` 헬퍼.
- `OllamaVisionOcr`: `<endpoint>/api/generate` 비스트리밍 호출,
`images: [base64]` 필드로 이미지 전달, 프롬프트는 언어 힌트
+ 화이트리스트 언어 목록 포함. 응답 prose 를 `OcrText.joined` 로,
prepared image 전체 영역 단일 region (confidence 1.0) 으로 wrap.
기본 모델 `gemma4:e4b`. endpoint 비어 있으면 `models.llm.endpoint`
로 fallback.
- 이미지 전처리: long-edge `config.image.ocr.max_pixels` (기본 1600,
256~4096 클램프) 초과 시 PNG 로 재인코딩 (image::imageops::resize,
Triangle filter). PNG 입력이 max 이내면 zero-copy passthrough.
- `apply_ocr` 는 OCR 성공 시 block.ocr 를 Some 으로 채우고
ProvenanceKind::OcrApplied 이벤트 추가. 실패 시 block.ocr 는
None 그대로 + provenance 미기록 (부분 상태 누출 금지).
- `kebab-config`: 새 `ImageCfg.ocr: OcrCfg` 블록 (enabled/engine/model
/endpoint/languages/max_pixels). `#[serde(default)]` 로 pre-P6
TOML 호환. `KEBAB_IMAGE_OCR_*` 환경변수 5종 추가.
## Spec deviation
원래 P6-2 spec 은 Tesseract 를 default OCR 엔진으로 지정했으나, dev /
CI 호스트에서 `libtesseract-dev` 시스템 패키지 설치를 피하려고
Ollama-vision 으로 default 를 교체. `OcrEngine` trait 추상화는 spec
그대로 보존 — Tesseract / Apple Vision / PaddleOCR 어댑터는 같은
trait 으로 추후 feature-gate 추가 가능. 자세한 내역은
`tasks/HOTFIXES.md` 2026-05-02 항목 참조.
Trust 측면: vision LM 은 hallucinate 가능. `OcrText.engine = "ollama-vision"`
필드로 consumer 가 엔진 별 신뢰 분기 가능.
## 테스트
- 신규 (`tests/ocr.rs`, 8 + 1 ignored):
- 200 happy → OcrText 디코딩 (joined / engine / engine_version /
region count / bbox / confidence)
- 빈 응답 → 빈 regions
- 5xx → Err with status + body 포함
- 200 error envelope → Err
- apply_ocr → block.ocr Some + Provenance OcrApplied 1건
- apply_ocr error → block.ocr None 유지 + events 미기록
- 4000×3000 PNG → max_pixels=1024 까지 다운스케일, aspect ratio 보존
- from_parts max_pixels 클램프
- opt-in `KEBAB_OCR_INTEGRATION=1` 통합 (실제 192.168.0.47 Ollama
`gemma4:e4b` 로 \"Hello World 2026\" 전사 검증 완료)
- 신규 (`src/ocr.rs` unit): truncate, build_prompt 언어/힌트 처리
- `kebab-config` 테스트 +3: defaults, env override, pre-P6 TOML 호환
전체: `cargo test -p kebab-parse-image` 28 pass + 1 ignored,
`cargo test -p kebab-config` 20 pass,
`cargo clippy --workspace --all-targets -- -D warnings` pass.
contract: docs/superpowers/specs/2026-04-27-kebab-final-form-design.md
sections: §3.4 ImageRefBlock.ocr, §3.7a OcrText / OcrRegion, §9.1 OCR
vs caption provenance.
This commit is contained in:
@@ -43,6 +43,60 @@ pub fn no_exif_png() -> Vec<u8> {
|
||||
buf.into_inner()
|
||||
}
|
||||
|
||||
/// 4000×3000 solid-blue PNG (long edge 4000) used to exercise the OCR
|
||||
/// adapter's downscale path. Solid-colour PNGs compress aggressively, so
|
||||
/// the on-disk size stays well under 1 MB despite the large dimensions.
|
||||
pub fn large_blue_4000x3000_png() -> Vec<u8> {
|
||||
let img: ImageBuffer<Rgb<u8>, _> =
|
||||
ImageBuffer::from_fn(4000, 3000, |_, _| Rgb([0, 0, 255]));
|
||||
let mut buf = Cursor::new(Vec::new());
|
||||
img.write_to(&mut buf, image::ImageFormat::Png)
|
||||
.expect("encoding 4000x3000 PNG must not fail");
|
||||
buf.into_inner()
|
||||
}
|
||||
|
||||
/// PNG with the literal text `"Hello World 2026"` rendered in black
|
||||
/// against a white background. Used by the opt-in
|
||||
/// `ocr_integration_real_ollama_transcribes_text` integration test —
|
||||
/// regular hermetic tests never call it.
|
||||
///
|
||||
/// Renders with a font shipped by `dejavu` (path under
|
||||
/// `/usr/share/fonts/truetype/dejavu/`) — common across most Linux dev
|
||||
/// boxes. Falls back to a tiny built-in glyph map if the font is
|
||||
/// missing so the helper compiles even without DejaVu installed.
|
||||
pub fn hello_world_png() -> Vec<u8> {
|
||||
use ab_glyph::{Font, FontRef, ScaleFont};
|
||||
|
||||
let mut img: ImageBuffer<Rgb<u8>, _> =
|
||||
ImageBuffer::from_fn(400, 100, |_, _| Rgb([255, 255, 255]));
|
||||
let font_bytes = std::fs::read("/usr/share/fonts/truetype/dejavu/DejaVuSans-Bold.ttf")
|
||||
.expect("DejaVu Sans Bold required for OCR integration fixture");
|
||||
let font = FontRef::try_from_slice(&font_bytes).expect("font parses");
|
||||
let scaled = font.as_scaled(40.0);
|
||||
let text = "Hello World 2026";
|
||||
let mut x = 10.0_f32;
|
||||
let y = 60.0_f32;
|
||||
for ch in text.chars() {
|
||||
let glyph = scaled.scaled_glyph(ch);
|
||||
if let Some(outlined) = scaled.outline_glyph(glyph.clone()) {
|
||||
let bb = outlined.px_bounds();
|
||||
outlined.draw(|gx, gy, c| {
|
||||
let px = (x + bb.min.x + gx as f32) as i32;
|
||||
let py = (y + bb.min.y + gy as f32) as i32;
|
||||
if px >= 0 && py >= 0 && (px as u32) < 400 && (py as u32) < 100 {
|
||||
let v = ((1.0 - c) * 255.0) as u8;
|
||||
img.put_pixel(px as u32, py as u32, Rgb([v, v, v]));
|
||||
}
|
||||
});
|
||||
}
|
||||
x += scaled.h_advance(scaled.glyph_id(ch));
|
||||
}
|
||||
let mut buf = Cursor::new(Vec::new());
|
||||
img.write_to(&mut buf, image::ImageFormat::Png)
|
||||
.expect("encoding hello-world PNG must not fail");
|
||||
buf.into_inner()
|
||||
}
|
||||
|
||||
/// JPEG with embedded EXIF APP1 segment carrying GPS + Make + Model +
|
||||
/// DateTimeOriginal + Orientation + Software. The base image is a 4×4
|
||||
/// solid white square — pixel content is irrelevant; the test cares about
|
||||
|
||||
Reference in New Issue
Block a user