kebab

Author	SHA1	Message	Date
altair823	aebf900b2f	docs(tasks): P7-3 pdf ingest wiring task spec P7-1 (`PdfTextExtractor`) + P7-2 (`PdfPageV1Chunker`) 의 라이브러리는 완성됐지만 `kebab-app::ingest` 가 `MediaType::Pdf` 를 dispatch 하지 않아 CLI 에서 PDF 가 보이지 않는 상태. P7-3 이 그 와이어링을 다룬다 — P6-4 의 image wiring 패턴과 평행. 핵심 결정 (spec 본문): - 새 private fn `ingest_one_pdf_asset` (P6-4 의 `ingest_one_image_asset` 와 평행). `ingest_one_asset` match 에 `MediaType::Pdf` arm 추가. - per-medium chunker 선택: PDF 는 `PdfPageV1Chunker` 하드코딩 (md 는 `MdHeadingV1Chunker` 그대로). `config.chunking.chunker_version` 은 PDF ingest 에서 무시 (deviation, HOTFIXES 추가 예정). - encrypted PDF / corrupt PDF → `errors+=1` + `IngestItem.error` 에 P7-1 의 `qpdf --decrypt` 안내 그대로 보존. - 빈/scanned candidate 페이지 → asset 인덱싱, 빈 페이지 0 chunk, P7-1 emit 한 `Provenance::Warning` 그대로 통과. 향후 OCR fallback 까지는 검색 불가 (out of scope). - determinism stress: extract → chunk 사이에 `now()` 추가 호출 금지 (P6-4 와 동일 invariant). - 11 통합 테스트 + smoke 업데이트 (별도 implementation PR). `tasks/INDEX.md` P7 components 2 → 3 반영. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>	2026-05-02 09:07:03 +00:00
altair823	7ee0ac9894	feat(kebab-chunk): P7-2 pdf-page-v1 chunker — page-aware splitting `PdfPageV1Chunker` 가 `kebab-parse-pdf` 가 emit 한 `CanonicalDocument` (블록당 한 페이지, 모두 `SourceSpan::Page`) 를 받아 페이지 경계를 절대 넘지 않는 `Chunk` 들을 생성. `chunker_version = "pdf-page-v1"`. 핵심 동작: - 페이지 텍스트가 `target_tokens × BYTES_PER_TOKEN` (= 3) 안이면 한 덩어리. 초과 시 `\n\n` (paragraph) 또는 sentence-end 구두점 + whitespace 경계를 segment 로 보고 greedy 누적, 기본 한 chunk 당 최소 한 segment. - 다음 chunk 의 prefix 에 `overlap_tokens × BYTES_PER_TOKEN` 만큼의 직전 꼬리를 prepend (char 단위, 이전 chunk 시작 너머로 backtrack 안 함). - 빈/공백-only 페이지는 0 chunk (페이지의 `Provenance::Warning` 으로 `kebab-parse-pdf` 단계에 이미 표시됨). - 비-PDF doc (Block::Paragraph 가 아니거나 SourceSpan 이 Page 아님) → 명시 에러. Spec deviation (HOTFIXES 2026-05-02 P7-2): - `chunk_id` 충돌 가드: 같은 페이지에서 여러 chunk 가 나오면 `block_ids` 가 모두 같아 §4.2 recipe 가 충돌. `id_for_chunk` 의 `policy_hash` 인풋을 per-chunk 로 `format!("{base}#c{char_start}")` 변형해 회피. recipe 자체는 불변. `Chunk.policy_hash` 필드는 base 유지. - `BYTES_PER_TOKEN = 3` (md-heading-v1 실제 코드와 일치). spec 본문은 "/ 4" 라고 했지만 그 자체가 md-heading-v1 의 실코드와 어긋나 있어 일관성 쪽을 택함. cross-chunker `policy_hash` 동일성 unit test 로 잠금. 테스트 (10개 신규): - chunker_version label, 3-page small, 1-page huge + overlap + chunk_id 유일성, empty page skip, whitespace-only skip, non-PDF error, cross-page boundary 절대 안 만들어짐, determinism (1000회), snapshot shape 안정, md-heading-v1 와 policy_hash 동일. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>	2026-05-02 08:51:44 +00:00
altair823	5a158d7343	feat(kebab-parse-pdf): P7-1 text PDF extractor — per-page CanonicalDocument `PdfTextExtractor`(MediaType::Pdf) lopdf 기반 per-page 텍스트 추출. 페이지마다 `Block::Paragraph` + `SourceSpan::Page { page, char_start, char_end }` emit. 본문이 비거나 추출 panic 인 페이지는 빈 paragraph + `Provenance::Warning` ("scanned candidate") 로 표시 — 이후 OCR fallback (별도 task) 의 입력. 핵심 동작: - `lopdf::Document::load_mem` + `is_encrypted()` → 암호화 PDF 는 명시 에러 (`qpdf --decrypt` 안내). - 페이지 단위 `extract_text(&[page])` 를 `catch_unwind` 로 감싸 malformed page panic 을 recoverable warning 으로 변환. - `/Info` dict 에서 Title/Producer/Creator best-effort 추출. UTF-16BE BOM prefixed 문자열도 디코드 (한국어 등 non-ASCII Title 정상 처리). - 9개 통합 테스트: 3-page emit, scanned-mixed warning, encrypted refuse, corrupt header error, page_count 메타, UTF-16BE Title, filename fallback, determinism, snapshot. `parser_version = "pdf-text-v1"`. Allowed deps: `lopdf 0.32` + `pdf-extract 0.7` (원본 spec 그대로). 본문 다국어 OCR fallback 은 §9.2 후속 task (out of scope). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>	2026-05-02 08:34:55 +00:00
altair823	f9714aa5cb	docs(rename): kb → kebab — README, tasks/, docs/, design doc, report 마지막 commit. 모든 .md 안의 `kb` 단어 일괄 갱신. - 19 개 crate 이름 (`kb-core`, `kb-app`, …) → `kebab-` (Rust 모듈 path 표기 `kb_` → `kebab_` 포함). - 미래 component (`kb-tui`, `kb-desktop`, `kb-asr-whisper`, `kb-ocr`, `kb-mcp`, `kb-vlm`, `kb-rerank`, `kb-vision-ocr`, `kb-index`, `kb-smoke`, `kb-architecture`) → `kebab-` (P6+ 가 시작될 때 같은 prefix 사용). - CLI 명령 예제: `kb ingest` / `kb search` / `kb ask` / `kb init` / `kb doctor` / `kb inspect` / `kb list` / `kb eval` → `kebab <verb>`. fenced code block + 인라인 backtick 모두. - XDG paths + env vars + binary 경로 (`target/release/kb` → `target/release/kebab`) 동기화. - design doc / 최초 보고서 / SMOKE / HOTFIXES / phase epic / task spec 모든 reference 통일. - task-decomposition.md 의 `git -c user.name=kb` 는 과거 git history 기록용 author 정보라 그대로 유지 (실제 git history 의 author 는 변경 불가). - `tasks/phase-5-evaluation.md` 의 `status: planned` → `completed` 도 같이 (P5-1 + P5-2 PR 머지 후 미반영분). ## 검증 - `grep -rEn "\bkb-[a-z]\|\bkb_[a-z]\|\.config/kb\b\|kb\.sqlite\|\bKB_[A-Z]" --include="*.md"` 0 hits (task-decomposition.md 의 git author 제외). - 모든 file path reference 살아있음 (renamed file 들 모두 새 path 로 update). 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>	2026-05-02 04:01:55 +00:00
kb	b999a12ab5	tasks: address PR #1 review - p3-3: SQLite-first/Lance-second + status marker (V003__embedding_status); drop "best-effort 2PC" misnomer - p4-3: replace print_stream FnMut closure with mpsc::Sender<String> (RagPipeline stays Send+Sync) - p4-3: tighten citation regex to strict [#<n>] only — reject [n]/prose/code-block false positives - p5-2: compare_runs across chunker_version is graceful (doc + span overlap fallback) with chunker_version_match audit field; --strict-chunker-version restores refusal - p7-1: per-page text via lopdf (pdf-extract has no per-page Rust API); use char count for spans - p8-1: explicit rubato (FftFixedIn) for 16 kHz mono resample; symphonia decode only - p9-5: drop cmd_read_pdf_page + pdfium native dep; cmd_read_file_bytes + frontend pdfjs; add traversal tests	2026-04-27 13:10:31 +00:00
kb	d96d9cc56c	tasks: add P7 component specs (pdf-extractor, pdf-chunker)	2026-04-27 12:07:56 +00:00

6 Commits