Merge pull request 'refactor: 척추 재작성 — 표면·동작·내부 단순화 (cuts + config 재편 + ingest 스파인)' (#214) from refactor/spine-cuts into main
Reviewed-on: #214
This commit was merged in pull request #214.
This commit is contained in:
@@ -4,7 +4,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co
|
||||
|
||||
## Project
|
||||
|
||||
Single-user local-first knowledge base + RAG. Rust 2024 workspace, 24 crates, single binary (`kebab`). All inference is local (Ollama + fastembed + whisper.cpp).
|
||||
Single-user local-first knowledge base + RAG. Rust 2024 workspace, 22 crates, single binary (`kebab`). All inference is local (Ollama + fastembed + whisper.cpp).
|
||||
|
||||
The repo's documentation is split by audience — don't duplicate across them:
|
||||
|
||||
@@ -31,7 +31,7 @@ The dev/test profile is already trimmed (`debug = "line-tables-only"`, `split-de
|
||||
|
||||
## The facade rule
|
||||
|
||||
`kebab-app` is the only crate UI binaries (`kebab-cli`, future `kebab-tui`, `kebab-desktop`) may touch. Every user-facing entry has a `*_with_config(cfg, …)` companion that takes an explicit `Config`:
|
||||
`kebab-app` is the only crate UI binaries (`kebab-cli`, future `kebab-desktop`) may touch. Every user-facing entry has a `*_with_config(cfg, …)` companion that takes an explicit `Config`:
|
||||
|
||||
- `kebab-cli` calls the `*_with_config` form so `--config <path>` is honored.
|
||||
- The bare `kebab_app::ingest(...)` / `search(...)` / `ask(...)` form re-loads `Config::load(None)` (XDG default) and silently bypasses any explicit path. Two regressions of exactly this shape are recorded in `tasks/HOTFIXES.md` (P3-5 + P4-3 follow-ups). When wiring a new CLI subcommand, always thread the `Config` through.
|
||||
@@ -54,7 +54,7 @@ Each task spec lists `Allowed dependencies` and `Forbidden dependencies` per des
|
||||
|
||||
- `kebab-core` MUST NOT depend on any other `kebab-*` crate. Domain types only.
|
||||
- `kebab-eval`'s `metrics` and `compare` modules MUST NOT import retrieval / embedding / LLM crates directly. The runner is allowed to use `kebab-app`'s facade (P5-1 inheritance — see deviations in that task spec).
|
||||
- UI crates (`kebab-cli`, `kebab-mcp`, `kebab-tui`, future `kebab-desktop`) MUST NOT import `kebab-store-*` / `kebab-llm-*` / `kebab-parse-*` directly — only `kebab-app`.
|
||||
- UI crates (`kebab-cli`, `kebab-mcp`, future `kebab-desktop`) MUST NOT import `kebab-store-*` / `kebab-llm-*` / `kebab-parse-*` directly — only `kebab-app`.
|
||||
|
||||
Read the relevant task spec's deps section before adding an import. New crates inherit the same boundary rules.
|
||||
|
||||
|
||||
933
Cargo.lock
generated
933
Cargo.lock
generated
File diff suppressed because it is too large
Load Diff
@@ -11,7 +11,6 @@ members = [
|
||||
"crates/kebab-search",
|
||||
"crates/kebab-embed",
|
||||
"crates/kebab-embed-local",
|
||||
"crates/kebab-embed-candle",
|
||||
"crates/kebab-embed-ollama",
|
||||
"crates/kebab-llm",
|
||||
"crates/kebab-llm-local",
|
||||
@@ -21,7 +20,6 @@ members = [
|
||||
"crates/kebab-eval",
|
||||
"crates/kebab-parse-image",
|
||||
"crates/kebab-parse-pdf",
|
||||
"crates/kebab-tui",
|
||||
"crates/kebab-mcp",
|
||||
"crates/kebab-parse-code",
|
||||
"crates/kebab-nli",
|
||||
@@ -93,7 +91,7 @@ struct_excessive_bools = "allow"
|
||||
naive_bytecount = "allow"
|
||||
# `#[ignore]` annotations on tests document via the test name + nearby comment.
|
||||
ignore_without_reason = "allow"
|
||||
# `format!` push patterns are a hot path for kebab-tui's progressive rendering;
|
||||
# `format!` push patterns are common in CLI output paths;
|
||||
# `write!` rewrite needs a verified-equal benchmark before swapping.
|
||||
format_push_string = "allow"
|
||||
# Builder-style `with_*` methods return `Self`; the existing `#[must_use]`
|
||||
@@ -140,9 +138,6 @@ rusqlite = { version = "0.32", features = ["bundled"] }
|
||||
globset = "0.4"
|
||||
tempfile = "3"
|
||||
proptest = "1"
|
||||
# p9-fb-19: LRU cache for `App::search` results. Bounded capacity
|
||||
# from `config.search.cache_capacity` (default 256, ~1.3 MB cap).
|
||||
lru = "0.12"
|
||||
lopdf = "0.32"
|
||||
# fastembed-rs ships ONNX runtime via the `ort-download-binaries` feature
|
||||
# in its default set (which also pulls `hf-hub` for first-run model
|
||||
|
||||
16
HANDOFF.md
16
HANDOFF.md
@@ -19,7 +19,7 @@ P0–P5 + P6 + P7 + P9-1/2/3/4 (Library / Search / Ask / Inspect) + P10 전체
|
||||
| **P6** | 이미지 ingestion (OCR + caption) | `kebab-parse-image` | P5 | ✅ 완료 (4/4 component, OCR/caption Ollama-vision) |
|
||||
| **P7** | PDF text + page citation + scanned OCR (v0.20.0 sub-item 1) | `kebab-parse-pdf` + `kebab-app::pdf_ocr_apply` | P5 + P6 | ✅ 완료 (3/3 component, page-level chunker + ingest wiring + post-extract OCR enrichment via qwen2.5vl:3b vision LLM) |
|
||||
| **P8** | 음성 transcription + timestamp citation | `kebab-parse-audio` | P5 | ⏸ 보류 (whisper-rs 시스템 dep brainstorm 필요) |
|
||||
| **P9** | TUI + desktop app | `kebab-tui`, `kebab-desktop` | P5 | 🟡 진행 (4/5 component — P9-1/2/3/4 완료 [Library / Search / Ask / Inspect], P9-5 desktop 예정 · 도그푸딩 피드백 **20/20 ✅**) |
|
||||
| **P9** | desktop app | `kebab-desktop` | P5 | ⏸ 보류 (TUI 제거됨 — CLI/MCP 로 UI 집약, P9-5 desktop 예정) |
|
||||
| **P10** | code ingest framework | `kebab-parse-code` | P5 | 🟡 진행 중 — 1A-1 ✅ (wire schema + parse-code skeleton + filter flags), 1A-2 ✅ (Rust AST chunker, `code-rust-ast-v1` — v0.7.0), 1B ✅ (Python/TS/JS AST chunkers — v0.8.0 이후), **1C-Go ✅ (Go AST chunker, `code-go-ast-v1` — v0.12.0)**, **1C-JavaKotlin ✅ (Java + Kotlin AST chunkers, `code-java-ast-v1` / `code-kotlin-ast-v1` — v0.13.0)**, **2 ✅ (Tier 2 resource-aware: yaml/k8s + dockerfile + manifest, `k8s-manifest-resource-v1` / `dockerfile-file-v1` / `manifest-file-v1` — v0.14.0)**, **3 ✅ (Tier 3 paragraph fallback: code-text-paragraph-v1 — v0.15.0)**, **1D ✅ (C + C++ AST chunkers, code-c-ast-v1 + code-cpp-ast-v1 — v0.16.0)** |
|
||||
|
||||
P0~P5 직렬. P6~P9 P5 이후 병렬 가능.
|
||||
@@ -30,7 +30,7 @@ P0~P5 직렬. P6~P9 P5 이후 병렬 가능.
|
||||
|
||||
## 머지 후 발견된 버그 / 결정 (요약)
|
||||
|
||||
- **candle 임베딩 백엔드 다변화** (2026-06-01, Track 1, v0.22.0): `provider = "candle"` opt-in 추가 — 같은 `multilingual-e5-large` 모델을 순수 Rust(candle)로 돌려 듀얼소켓 NUMA 서버의 onnxruntime 48-스레드 double-free 를 회피. `[models.embedding].num_threads`(+env `KEBAB_EMBED_THREADS`)로 CPU 스레드 캡. fastembed default 동작·벡터 불변, `embedding_version` 유지(재색인 0). Phase 0 스파이크 패리티 cosine 1.000000. 상세 HOTFIXES 동일 일자.
|
||||
- **candle 임베딩 백엔드 제거** (2026-06-24, spine-phase0): `provider = "candle"` 제거 — fastembed(기본) + ollama 두 경로로 충분, candle 은 미사용. `num_threads` 필드는 TOML 하위호환 위해 유지(레거시, 무시). 상세: spine-cuts 계획.
|
||||
- **config 마이그레이션** (2026-05-31, PR #198): `kebab config migrate` 추가 — 기존 config.toml 에 빠진 섹션을 주석과 함께 채우고 deprecated 정리(멱등·`.bak`·dry-run, 값/주석 보존). `schema_version` 1→2, `init` 도 섹션 주석 포함, doctor 에 `config_migration` 체크. 상세 HOTFIXES 동일 일자.
|
||||
|
||||
머지 후 발견된 모든 deviation / hotfix 의 dated 로그는 [tasks/HOTFIXES.md](tasks/HOTFIXES.md). 본 요약은 \"누군가가 인수받을 때 알아두면 시간을 많이 절약하는\" 항목만:
|
||||
@@ -40,7 +40,7 @@ P0~P5 직렬. P6~P9 P5 이후 병렬 가능.
|
||||
- **2026-06-04 PP-OCRv5 ONNX Rust 네이티브 OCR** — v0.27.0. `[image.ocr] engine = "paddle-onnx"` 로 PP-OCRv5(검출+인식) ONNX 를 in-process(`ort` =2.0.0-rc.9) 실행 — Python 런타임/원격 호출 없이 큰 페이지 CPU <4초(Ollama vision ~50초 대비). default 는 여전히 `"ollama-vision"`. 후처리(min-area rect/unclip)는 pure-Rust. **함정**: unclip 은 corner 를 centroid 에서 방사 확장하면 안 되고 edge 별 polygon offset 이어야 함(방사 확장 시 wide/short 텍스트 박스 높이가 안 커져 글자 윗부분 잘림 → ㄷ→ㄴ, e2e CER 0.26). 수정 후 CER 0.005. 모델 ONNX 는 `crates/kebab-parse-image/assets/paddleocr-onnx/`(LFS). 자세한 내용: `tasks/HOTFIXES.md` (2026-06-04 PP-OCRv5 ONNX), spec/plan `docs/superpowers/{specs,plans}/2026-06-04-rust-native-ocr-*.md`.
|
||||
- **2026-06-03 ingest 설정 변경 자동 재색인** — v0.26.2. ingest 산출에 영향 주는 설정(청킹/이미지 OCR·caption/pdf.ocr/`[ingest.code]`)을 변경하면 `--force-reingest` 없이 영향 자산만 자동 재색인. 그 설정들의 결정적 서명(`ingest_config_signature`)을 effective parser_version(skip 비교 + 저장 doc 필드 양쪽)에 폴딩 → 다음 ingest 비교가 mismatch. 비산출 설정(search/rag/ui/log + max_pixels/languages/timeout)은 제외(과도 무효화 회피), doc_id 는 base 로 안정 유지. **업그레이드 후 첫 ingest 는 전 자산 1회 재색인**(저장된 상수 parser_version ≠ 새 composite; embedding 은 V012 캐시 히트). 결과 포맷·CLI·wire 불변(내부 skip 판정 정정). 자세한 내용: `tasks/HOTFIXES.md` (2026-06-03 ingest 설정 변경 자동 재색인), spec/plan `docs/superpowers/{specs,plans}/2026-06-03-*invalidation*.md`.
|
||||
- **2026-06-03 ingest 진행 로그 개선** — v0.26.1. 이미지/PDF + OCR/caption on 볼트 ingest 가 "멈춘 듯" 보이던 문제 해소: TTY 진행바에 현재 파일명 + 느린 phase(ocr/caption/embed)+모델명 + 경과초 `(Ns)` heartbeat, 종료 시 최장 소요 파일 top-5 요약. 신규 wire `asset_phase{idx,total,phase,model}` + `asset_timings.ocr_ms`/`caption_ms`(additive, `ingest_progress.v1` 유지, serde default 0). 이미지·PDF 경로도 `asset_timings` emit(이전 markdown 만). 기본 동작 불변. 자세한 내용: `tasks/HOTFIXES.md` (2026-06-03 ingest 진행 로그), spec/plan `docs/superpowers/{specs,plans}/2026-06-03-ingest-log-improve-*.md`.
|
||||
- **2026-06-03 arctic-embed-l-v2.0 임베더 통합** — v0.26.0. 별칭 제거 후 설명형 query recall 보강(측정 recall@10 130/132, e5 +7). `kebab-embed-candle` 모델 레지스트리화(e5 mean + `snowflake-arctic-embed-l-v2.0` CLS, 모델별 pooling/prefix) + 신규 `kebab-embed-ollama`(`provider="ollama"`, `/api/embed`). config `endpoint: Option<String>` 추가. 기본 e5 유지(opt-in), arctic 전환은 embedding_version cascade → 재색인. candle↔Ollama cosine>0.99 게이트로 pooling/prefix 정확성 고정(`#[ignore]`). 자세한 내용: `tasks/HOTFIXES.md` (2026-06-03 arctic), spec `docs/superpowers/specs/2026-06-03-arctic-embedder-spec.md`.
|
||||
- **2026-06-03 arctic-embed-l-v2.0 임베더 통합** — v0.26.0. 별칭 제거 후 설명형 query recall 보강(측정 recall@10 130/132, e5 +7). 신규 `kebab-embed-ollama`(`provider="ollama"`, `/api/embed`). config `endpoint: Option<String>` 추가. 기본 e5 유지(opt-in), arctic 전환은 embedding_version cascade → 재색인. 자세한 내용: `tasks/HOTFIXES.md` (2026-06-03 arctic), spec `docs/superpowers/specs/2026-06-03-arctic-embedder-spec.md`.
|
||||
- **2026-06-03 doc-side expansion(별칭) 기능 완전 제거** — v0.25.0. 아래 2026-05-31 항목의 색인-시 청크당 LLM 별칭 생성 + 별칭 검색 채널을 **전부 제거**(ROI 음수: cross-lingual 은 e5-large 단독으로 충분, 기여는 설명형 +2 그룹뿐인데 대가가 청크당 색인-시 LLM). `Chunk.aliases`/`expansion.rs`/`IngestExpansionCfg`/alias lexical arm/`expansion_progress` wire kind 제거, 신규 마이그레이션 **V013** 이 `chunk_aliases_fts`+`chunks.aliases` DROP. 별칭 default-off 였어 사용자 체감 0, 기존 KB 도 재색인 불요(잔존 별칭 벡터는 `strip_alias_suffix` graceful 매핑/`reset` 정리). `AssetTimings.expansion_ms` 는 wire 호환 위해 값 0 으로 유지. 자세한 내용: `tasks/HOTFIXES.md` (2026-06-03), spec `docs/superpowers/specs/2026-06-03-remove-doc-expansion-spec.md`.
|
||||
- **2026-05-31 Phase 2 doc-side expansion 별칭(개별 dense 벡터) + 파생물 캐시(V012)** — v0.21.0 cut. 색인 시 LLM 이 청크별 별칭("같은 의미 다른 표현")을 생성, 줄별 **개별 dense 벡터**(sentinel `{chunk}#alias#N`)로 색인 (묶음 1벡터는 평균화 희석으로 회귀 → 폐기) + boilerplate 청크 skip. `[ingest.expansion]` default off. 측정(나무위키 ~1000 문서 CS corpus): 변형 일관성 14/18 → **16/18**, spread 0.222→0.111, 대조군 false-positive 별칭 무죄. 비용 병목(별칭 18문서 2.5h)은 **파생물 캐시(V012, 청크 내용 해시 키)**로 해소 — 정답 3개 cold 1879s → warm 13s **≈ 145배**, embedding+별칭 LLM 캐싱, version_key cascade 정합. search/ask 가 `kebab.sqlite`+`lancedb` 만으로 동작 → 외부 서버 색인 후 DB 만 복사하는 이식 워크플로 가능. **결정/known limitation**: grounded/refusal 판정이 부분 인용을 grounded 로 오분류(정직한 거부가 false-positive 로 집계) — 별도 개선 후보. stack·svm 설명형 2개 잔존. 자세한 내용: `tasks/HOTFIXES.md` (2026-05-31), 측정: `docs/superpowers/handoffs/2026-05-31-namu-wiki-alias-cache-study.md`.
|
||||
- **2026-05-29 v0.20.2 dogfood findings + 검색 품질 baseline** — 8-finding 라운드 완료. (1) Ask 응답언어: rag-v3 default (질문 언어 = 답변 언어). (2) eval `--config` facade 패치 로 dogfood KB 직접 eval 가능. (3) 검색 품질 baseline — hybrid hit@3=1.0 / MRR=0.833, lexical hit@3=1.0 / MRR=0.7 (golden 10 query). **O-2 known limitation**: 소형 모델(gemma4:e4b) refusal 메시지의 query 언어 불일치 가능 — 판정은 정상, 표시 문구만 해당. 자세한 내용: `tasks/HOTFIXES.md` (2026-05-29).
|
||||
@@ -78,27 +78,17 @@ P0~P5 직렬. P6~P9 P5 이후 병렬 가능.
|
||||
- **2026-05-02 P9 도그푸딩 후속 (p9-fb-16)** — TUI Ask conversation UI. `AskState` 가 `turns: Vec<Turn>` + `current_question` + `conversation_id` + `last_answer` 로 재설계. answer area 가 transcript (`Q1/A1`, `Q2/A2`, ...) 로 갈음, 매 Enter 가 이전 turns 를 `history` 로 worker 에 전달 (`ask_with_history`). conversation_id 는 첫 submit 시 timestamp-based 자동 생성 (`conv_<unix_nanos_hex>`). `Ctrl-L` 가 turns + conversation_id 초기화 (in-flight worker 는 그대로 finish, 결과는 새 conversation 의 stale turn 으로 silently 폐기). spec: `tasks/p9/p9-fb-16-tui-ask-conversation.md`.
|
||||
- **2026-05-03 P9 도그푸딩 후속 (p9-fb-20)** — `kebab ask` 의 CLI citation block. 답변 출력 후 `근거:` 절 — `[N] <full path>#<fragment> (score=<s>)` 한 줄씩. `--show-citations` (default ON) / `--hide-citations` (pipe 시 답변 본문만) flag. `--json` 모드는 무영향 (citations 가 항상 wire payload 에 포함). spec p9-fb-20 의 \"TUI citation pane + jump\" 부분은 P9-3 의 기존 `render_citations_or_explain` 가 일부 cover — 추가 기능 (turn 별 fold + Enter/o jump + i inspect) 은 후속 task 로 미룸 (사용자 도그푸딩 priority 5위 의 핵심 = full path 가독성 = CLI block 으로 충족). spec: `tasks/p9/p9-fb-20-citation-surface.md`.
|
||||
- **2026-05-03 P9 도그푸딩 후속 (p9-fb-07)** — Markdown title fallback chain. `kebab-normalize::derive_title(frontmatter_title, &[Block], file_stem)` — 1) frontmatter title → 2) 첫 H1 → 3) 첫 H2 → 4) 첫 paragraph 80 chars → 5) 파일 stem (모든 단계 NFC 정규화, 빈 문자열 절대 반환 안 함, 마지막 sentinel `"untitled"`). `build_canonical_document` 가 lift 후 helper 호출. parser_version 상수 `pulldown-cmark-0.x` → `md-frontmatter-v2` bump — 기존 doc 은 `doc_id` 가 갱신되므로 다음 ingest 가 자동 재처리 (idempotent upsert, design §9 cascade). spec: `tasks/p9/p9-fb-07-md-title-fallback.md`.
|
||||
- **2026-05-03 P9 도그푸딩 후속 (p9-fb-09)** — TUI external editor return restore. Search `g` 키 (citation jump) 후 TUI 화면이 깨지는 버그 수정. `kebab-tui::editor::with_external_program(&mut TuiTerminal, Command)` helper 가 suspend (LeaveAlternateScreen + Show cursor + disable_raw_mode) → spawn → restore (enable_raw_mode + EnterAlternateScreen + Hide cursor + `terminal.clear()`) 시퀀스를 RAII guard 로 atomic 하게 묶음. `App.pending_editor: Option<EditorRequest>` + `App.force_redraw: bool` 추가 — 키 핸들러는 EditorRequest enqueue 만, 실제 spawn 은 run loop 가 `TuiTerminal` 핸들 들고 처리. 후속 task (p9-fb-20 의 citation jump 등) 가 같은 helper 위에 build. spec: `tasks/p9/p9-fb-09-tui-editor-restore.md`.
|
||||
- **2026-05-03 P9 도그푸딩 후속 (p9-fb-14)** — TUI color theme module. `kebab-tui::theme::{Theme, Role, Palette}` 신규 — 16 개 Role (BorderActive/Title/Path/ModeLexical/ModeVector/ModeHybrid/Selected/Hint/Heading/Warning/Error/Success/CitationMarker/Bullet/Body/BorderInactive) 을 dark + light 두 팔레트가 exhaustive match 로 매핑. 모든 Pane (library/search/ask/inspect/run/error_popup) 의 inline `Style::default().fg(Color::*)` 호출이 `theme.style(Role::X)` 로 격리됨. `Config.ui.theme: String` (default `"dark"`) 신규. `App.theme: Theme` 가 `App::new` 에서 `Theme::from_name(&config.ui.theme)` 로 build — 알 수 없는 값은 dark fallback (config 가 typo 로 죽지 않음). `T` 키 runtime toggle 은 mode machine (p9-fb-12) 미진행이라 skip — config 만으로 결정. p9-fb-11 (ask markdown render) 의 Theme 의존성 unblock. spec: `tasks/p9/p9-fb-14-tui-color-theme.md`.
|
||||
- **2026-05-03 P9 도그푸딩 후속 (p9-fb-11)** — TUI Ask 답변 본문 markdown 렌더. `kebab-tui::markdown::render(text, &Theme) -> Vec<Line<'static>>` 신규 — `pulldown-cmark = "0.13"` 위에서 inline (bold/italic/strikethrough/inline code/link)·block (heading H1-H6, ordered/unordered list with nesting, fenced code block, table, blockquote `▎`, horizontal rule) 변환. heading H1/H2 = `Role::Heading`, H3+ = `Role::Title`, link = `Role::CitationMarker + UNDERLINE`, code = `Role::Hint`. ask `push_turn_lines` 가 grounded 답변에서만 markdown 렌더; refusal (`Role::Warning`) / streaming (`Role::Hint`) 은 raw 로 두어 role color 시그널 보존. CLI `kebab ask` 출력은 raw markdown 그대로 (terminal 호환성). 매 frame 재 parse — pulldown 토크나이저가 µs/KB 라 비용 무시. spec: `tasks/p9/p9-fb-11-ask-markdown-render.md`.
|
||||
- **2026-05-03 P9 도그푸딩 후속 (p9-fb-08)** — TUI search async worker + generation counter. 기존 200ms debounce 후 `kebab_app::search_with_config` 동기 호출이 vector/hybrid 모드 50-200ms 동안 UI freeze 시키던 문제 해소. `SearchState` 에 `generation: u64` + `worker_thread: Option<JoinHandle>` + `worker_rx: Option<Receiver<SearchWorkerMessage>>` 신규. `fire_search` 가 spawn 만 하고 즉시 return — worker 가 별 thread 에서 검색 후 `(generation, Result)` 를 channel 로 post. run loop 가 매 tick `poll_worker` 로 try_recv, generation 일치 시 hits 적용 / 불일치 시 silently 폐기 (사용자가 더 빠르게 타이핑하면 stale 결과 자동 drop). debounce_due 가 `searching && last_query == 현 input` 케이스 추가 skip — in-flight worker 의 결과 기다리는 동안 동일 query 재 spawn 안 함. spec: `tasks/p9/p9-fb-08-search-debounce.md`.
|
||||
- **2026-05-03 P9 도그푸딩 후속 (p9-fb-05)** — `workspace.root` path policy 명확화. `kebab_config::expand_path_with_base(raw, data_dir, base_dir) -> PathBuf` 신규 — 기존 `expand_path` (tilde + env 만) 위에 relative path resolution 추가, 절대/`~`/`${VAR}` 입력은 base_dir 무시. `Config.source_dir: Option<PathBuf>` 필드 (`#[serde(skip)]`) 신규 — `from_file` / `load` 가 `path.parent()` 로 stamp. `Config::resolve_workspace_root()` helper 가 `expand_path_with_base(&workspace.root, "", source_dir.unwrap_or(cwd))` 호출. kebab-app + kebab-source-fs 의 모든 `workspace.root` 사용 사이트가 `cfg.resolve_workspace_root()` 로 통일 — kebab-source-fs 의 fork 된 `expand_tilde` 헬퍼는 제거 (kebab-app 의 `storage.data_dir` 한 곳만 남음, P+ 통일 caveat). `kebab init` 가 생성하는 `config.toml` 위에 path policy 안내 헤더 코멘트 자동 prepend (절대/tilde/env/상대 + 상대 base = config dir). spec: `tasks/p9/p9-fb-05-config-path-policy.md`.
|
||||
- **2026-05-03 P9 도그푸딩 후속 (p9-fb-19)** — In-process LRU search cache + `corpus_revision` 카운터. SQLite V004 migration 으로 `kv (key TEXT PK, value TEXT)` 테이블 + `corpus_revision = '0'` seed. `SqliteStore::corpus_revision()` / `bump_corpus_revision()` 메서드 (`UPDATE ... CAST AS INTEGER + 1` 으로 atomic). `kebab-app::ingest_with_config_cancellable` 가 `new + updated > 0` 시 bump — no-op reingest 는 cache 보존. `App.search_cache: Option<Mutex<LruCache<SearchCacheKey, Vec<SearchHit>>>>` (capacity from `config.search.cache_capacity`, default 256, 0 = 비활성). `SearchCacheKey` = `query_norm` (NFKC + trim + lowercase) + `mode` + `k` + `snippet_chars` + `embedding_version` + `chunker_version` + `corpus_revision` snapshot. `App::search` 가 lookup → miss 시 `search_uncached` → put. `search_uncached_with_config` facade 추가, CLI `kebab search --no-cache` 로 bypass (디버깅용). frozen design §9 versioning 표에 `corpus_revision` row 추가. spec: `tasks/p9/p9-fb-19-search-cache.md`.
|
||||
- **2026-05-03 P9 도그푸딩 후속 (p9-fb-17)** — Multi-turn chat session 영속화 (storage 만 — UI 는 p9-fb-18). SQLite V005 migration (spec 의 V004 가 p9-fb-19 의 kv 와 충돌해서 V005 로 시프트, HOTFIXES) 으로 `chat_sessions` (session_id PK + created_at + updated_at + title + config_snapshot_json) + `chat_turns` (turn_id PK + session_id FK ON DELETE CASCADE + turn_index + question + answer + citations_json + created_at, UNIQUE(session_id, turn_index)) + `idx_chat_turns_session` 추가. `kebab_core::ChatSessionRepo` trait 6 메서드 (create_session / get_session / list_sessions / delete_session / append_turn / list_turns) + `kebab_core::{ChatSessionRow, ChatTurnRow}` 신규 export. `kebab-store-sqlite::SqliteStore` impl (별 `chat_sessions.rs` 모듈) — append_turn 이 insert + parent updated_at bump 을 같은 conn 에서 처리. frozen design §5 storage 에 §5.7a chat_sessions/turns 절 신설. spec: `tasks/p9/p9-fb-17-chat-session-storage.md`. unblocks p9-fb-18 (CLI session/repl).
|
||||
- **2026-05-03 P9 도그푸딩 후속 (p9-fb-18)** — CLI `kebab ask --session <id>` (multi-turn). p9-fb-17 의 ChatSessionRepo 위에 `kebab-app::App::ask_with_session(session_id, query, opts) -> Answer` 메서드. 첫 호출 시 자동으로 `chat_sessions` row 생성 (title = 첫 question NFC trim 40 chars), 이후 호출은 `list_turns` 로 prior history 받아 `RagPipeline::ask_with_history` 호출 + 새 turn append. `App` 의 helper: `first_question_title(question)` (NFC + trim + 40 char cap, fallback `"untitled"`) + `blake3_truncate(input)` (32-hex `turn_id` 생성). facade `kebab_app::ask_with_session_with_config` + CLI `--session <id>` flag 추가. `--repl` 은 spec 명시 사항이지만 stdin loop fixture 부담 으로 후속 task 로 deferral (out of scope per HANDOFF). spec: `tasks/p9/p9-fb-18-cli-ask-session-repl.md`.
|
||||
- **2026-05-03 P9 도그푸딩 후속 (p9-fb-12 partial)** — TUI vim-style mode machine (절반 ship — heuristic 제거는 follow-up). `kebab_tui::Mode::{Normal, Insert}` enum + `Mode::auto_for(pane)` (Library/Inspect/Jobs → Normal, Search/Ask → Insert) + `Mode::label()` (`"-- NORMAL --"` / `"-- INSERT --"`) + `App.mode: Mode` field. run loop `mode_intercept(app, key)` 가 dispatch 전 intercept — Insert 에서 `Esc` → Normal (어디서나), Normal 에서 `i` → Insert (Library/Inspect/Jobs 만, Search/Ask 는 자동 Insert 라 `i` 가 typed char). 헤더 우측에 mode label colored (Insert = Role::Success green, Normal = Role::Heading cyan+bold). pane 전환 시 `app.mode = Mode::auto_for(p)` 자동 flip. **Deferred (HOTFIXES entry)**: `is_typing_mod` (search) + input-empty heuristic (ask) 는 후속 PR 에서 mode-authoritative 로 교체 — 현재는 user-visible signal (label + auto flip + i/Esc) 만 ship, 키 dispatch 는 heuristic 유지. spec status `in_progress` (not `completed`). spec: `tasks/p9/p9-fb-12-tui-mode-machine.md`.
|
||||
- **2026-05-03 P9 도그푸딩 후속 (p9-fb-12 follow-up)** — heuristic 제거 (partial PR 의 deferred 부분 finalize). `search::is_typing_mod` (CTRL/ALT chord filter) 함수 삭제 + `ask::handle_key_ask` 의 input-empty heuristic 삭제. 새 dispatch: `search::handle_key_search` 의 `i` (chunk inspect) / `g` (editor jump) pre-pass 가 `state.mode == Mode::Normal` 일 때만 fire (Insert 에서는 typed char). main match 의 `j`/`k`/Char(c) 가 `state.mode` 로 분기 (Normal → 선택 이동, Insert → input.push). `ask::handle_key_ask` 의 `e`/`j`/`k` 도 동일 패턴 — Normal 에서 toggle/scroll, Insert 에서 input typing. 테스트 fixture (`tests/search.rs::fresh_app`, `tests/ask.rs::fresh_app`) 가 `app.mode = Mode::auto_for(focus)` 로 run-loop 동작 mirror. 기존 nav 테스트 (j_k_move, g_key_enqueues, e_toggles) 는 explicit `app.mode = Mode::Normal` 추가, 신규 4 테스트 (j_in_insert_types / arbitrary_char_in_normal_noop / e_types_in_insert / jk_scroll-in-normal-type-in-insert) 가 mode-authoritative 동작 pin. spec status `in_progress` → `completed`. spec: `tasks/p9/p9-fb-12-tui-mode-machine.md`.
|
||||
- **2026-05-03 P9 도그푸딩 후속 (p9-fb-10 partial)** — TUI CJK rendering helpers. `kebab-tui::input::{display_width, truncate_to_display_width}` 신규 — `unicode-width` 위에서 column-단위 width 계산 (ASCII=1, Hangul/CJK/fullwidth=2, combining=0) + char-boundary 안전 truncate (wide char 를 split 없이 keep-or-omit, ellipsis 1 col). library.rs 의 중복 `truncate_to_display_width` private fn 제거 — 단일 source. 9 unit tests (ASCII / Hangul / Japanese / mixed / truncate fits·overflow·zero-cols·wide-char-boundary / `String::pop` char-aware sanity) + 1 integration render test (Korean + Japanese fixture, TestBackend 80×20, 한글/일본어 글자가 frame 에 살아남음 확인). spec 의 `InputBuffer` struct (cursor 가 column 단위 wide-char width 추적) 도입은 follow-up — Ask/Search/Editor pane 의 String + cursor 일괄 마이그레이션이 회귀 표면이 커서 helper 만 먼저 머지. backspace 는 모든 pane 이 이미 `String::pop()` 사용 (char-aware) → byte-boundary 안전성 helper 없이도 확보. crossterm 0.28 이 native IME composing 미노출 — preedit handling out of scope. spec status `planned` → `in_progress`. spec: `tasks/p9/p9-fb-10-tui-cjk-input.md`.
|
||||
- **2026-05-04 P9 post-도그푸딩 (p9-fb-23)** — Incremental ingest. 사용자 도그푸딩 피드백: 변하지 않은 문서는 다시 ingest 하지 않기. blake3 checksum + parser_version + chunker_version + embedding_version 4개 input 이 모두 일치할 때 parse/chunk/embed/vector upsert 모두 회피. SQLite V006 마이그레이션 — `documents` 에 `last_chunker_version` + `last_embedding_version` 컬럼 추가. 신규 `IngestItemKind::Unchanged` variant + `IngestReport.unchanged` + `AggregateCounts.unchanged` (wire schema additive). `IngestOpts { progress, cancel, force_reingest }` struct 도입 — `AskOpts` 패턴. `--force-reingest` CLI flag 로 skip 우회. 비용 dominator (fastembed) 가 변경된 / 새 doc 에만 발생. spec: `tasks/p9/p9-fb-23-incremental-ingest.md`. HOTFIXES `2026-05-04 — p9-fb-23` 항목이 version cascade 명시 동작의 source of truth.
|
||||
- **2026-05-05 P9 post-도그푸딩 (p9-fb-25)** — Config 의 `workspace.include` 필드 제거 + 지원 형식 가시성. 사용자 도그푸딩 피드백: include + exclude 동시 존재가 case 4 (둘 다 매치 안 함) 의미 모호 + 어차피 처리 가능 형식 (md / png / jpg / pdf) 이 정해져 있으니 명시 필요. `WorkspaceCfg.include` 제거 (옛 config 의 `include = [...]` 은 silently 무시 + 단발 deprecation warning). `IngestItem.warnings` 가 Skipped 시 사유 (`"unsupported media type: .docx"` 등) 채움. `IngestReport.skipped_by_extension: BTreeMap<String, u32>` 신규 (additive wire — release 트리거 안 됨). CLI / TUI summary 에 breakdown 표시 (`"5 skipped: 3 docx, 1 txt, 1 epub"`). README + `kebab init` 헤더 주석에 지원 형식 명시. spec: `tasks/p9/p9-fb-25-config-include-removal.md`. HOTFIXES `2026-05-05 — p9-fb-25` 가 source of truth.
|
||||
- **2026-05-04 P9 post-도그푸딩 (p9-fb-24)** — TUI status/key bar + Library 컬럼 헤더 + Ask/Inspect PgUp/PgDn. 사용자 도그푸딩 3 건 (Library 컬럼 의미 부재, 페이지 스크롤 키 부재, 상태바 + 버전 정보 항상 노출 요청) 을 단일 PR 로 통합. bottom 영역을 status bar (1 row, version + pane + docs + dynamic state) + key hint bar (1 row, 기존 `footer_hints` 그대로) 두 줄로 분할; 기존 ingest progress dedicated row 는 status bar 의 dynamic slot 에 흡수 (priority cascade: streaming → searching → indexing → idle). Library `List` 위에 `format_doc_header` 행 + Layout 분할로 헤더 표시 (TITLE / TAGS / UPDATED / CHUNKS, display-width 정렬). `kebab-tui::pager::PAGE_STEP = 10` 신규 — Ask 의 PgUp/PgDn 추가 + Inspect 의 기존 +/-10 hardcode 가 같은 상수 참조로 통일. Ask 의 page-scroll 은 `j`/`k` 와 동일하게 `follow_tail = false` 로 freeze. spec: `tasks/p9/p9-fb-24-tui-affordances.md`. HOTFIXES `2026-05-04 — p9-fb-24` 항목이 footer 단행 row (p9-fb-13) + ingest dedicated row (p9-fb-03) 와의 layout 충돌의 source of truth.
|
||||
- **2026-05-04 P9 post-도그푸딩 (p9-fb-22)** — TUI 입력 cursor mid-string 편집 + Ask follow-tail auto-scroll. Gitea #94 (입력 후 커서 이동 안 됨) + #95 (새 응답 자동 스크롤 안 됨) 두 건. `InputBuffer` 의 cursor 모델을 byte-position 기반으로 재구성 — cursor 가 끝일 때 기존 append 동작과 backwards-compatible, mid-string 일 때는 `←/→/Home/End/Delete` 로 편집. `AskState` 에 `follow_tail: bool` (default true). `Paragraph::line_count(width)` (ratatui `unstable-rendered-line-info` feature 활성화) 로 매 프레임 wrapped row 수 계산해 follow-tail 시 scroll 을 bottom 에 pin. `j`/`k` 가 follow-tail 끄고 `Shift-G` 가 다시 켬. 12 신규 InputBuffer unit + 6 신규 Ask integration. spec: `tasks/p9/p9-fb-22-tui-cursor-and-autoscroll.md`. HOTFIXES 항목 `2026-05-04` 가 live cursor 모델 source of truth.
|
||||
- **2026-05-03 P9 post-도그푸딩 (p9-fb-21)** — `i` 가 universal Normal→Insert toggle (모든 pane). 이전 mode_intercept 는 Library/Inspect/Jobs 만 `i` intercept 였고 Search/Ask 는 fall-through (자동 INSERT 가정). 사용자가 Esc 로 NORMAL 로 빠진 후 Insert 복귀 키 없어 dead-end → 도그푸딩에서 보고됨. mode_intercept 의 `(Char('i'), Normal, _)` arm 이 pane 무관 모두 INSERT flip. Search 의 chunk inspect 키 `i`→`o` rebind (vim "open") 으로 충돌 해소. footer hint 모든 (pane, mode, filter) 조합 첫 fragment = `F1 도움말` (cheatsheet binding discoverability). Search/Ask Normal hint 에 `i 입력모드` fragment 추가. cheatsheet popup Global/Search/Ask section 갱신. 6 신규 unit + 3 기존 갱신. spec: `tasks/p9/p9-fb-21-tui-insert-key-discoverability.md` (status `completed` 직접). HOTFIXES 항목이 Search `i`→`o` rebind 의 source of truth.
|
||||
- **2026-05-03 P9 도그푸딩 후속 (p9-fb-10 follow-up)** — InputBuffer struct + 모든 text-input pane 마이그레이션 + cursor column 정렬. `kebab-tui::input::InputBuffer { content, cursor_col }` 신규 — `push_char` / `pop_char` / `clear` / `take` 가 wide-char 단위로 cursor_col 진행 (ASCII=1, Hangul/CJK=2, combining=0). `SearchState.input` / `AskState.input` / `FilterEdit.{tags_buf, lang_buf}` 가 InputBuffer 로 교체. render 단계에서 `f.set_cursor_position(...)` 가 `block.inner(area)` 기반 prompt 폭 + cursor_col 으로 caret 을 정확한 column 에 배치 (right-edge clamp). ratatui 0.28 의 cursor visibility 는 `cursor_position` Some/None 으로 자동 결정 — Search/Ask/Filter 가 `Some` 이라 caret 보임, Library/Inspect 는 `None` 이라 hidden. Korean lexical 검색은 `crates/kebab-app/tests/search_korean.rs` 에서 ingest → search → 결과 한 건 이상 + Korean 파일 stem 매칭 assert 로 회귀 핀. `lexical_query` test helper 가 `crates/kebab-app/tests/common/mod.rs` 로 promotion. spec status `in_progress` → `completed`. spec: `tasks/p9/p9-fb-10-tui-cjk-input.md`.
|
||||
- **2026-05-07 P9 post-도그푸딩 (p9-fb-27)** — `kebab schema [--json]` introspection 명령 + `error.v1` wire 도입. 정적 (wire schemas / capabilities / models) + 동적 (stats) 한 번에. `--json` 모드에서 fatal error 가 stderr ndjson 으로 emit (비 `--json` 은 기존 stderr text 유지). exit code 0/1/2/3 unchanged — `error.v1.code` 가 fine-grained 분기. fb-30 MCP `initialize` capability matrix 의 prerequisite. spec: `tasks/p9/p9-fb-27-introspection-and-error-wire.md`. design: `docs/superpowers/specs/2026-05-07-p9-fb-27-introspection-and-error-wire-design.md`.
|
||||
- **2026-05-03 P9 도그푸딩 피드백 20/20 ✅** — `tasks/p9/p9-fb-01..20` 모든 spec status `completed`. 사용자가 `kebab` 직접 돌려서 수집한 UX 잡음 (ingest 진행 표시 부재, mode 혼란, CJK column drift, multi-turn 부재, citation 부재 등) 이 모두 코드 또는 spec-acknowledged-deferred 형태로 해소. 도그푸딩 사이클 한 바퀴 완성 — P9-5 desktop tauri 와 별개로 TUI/CLI 사용자 경험 측면은 한 단계 안정화. P9 phase row 는 P9-5 미진행이라 🟡 유지.
|
||||
- **2026-05-03 P9 도그푸딩 후속 (p9-fb-13 follow-up)** — verb-form hint line 재구성. `pub fn footer_hints(focus: Pane, mode: Mode, filter_open: bool) -> &'static str` 신규 (run.rs). 한국어 동사구 (`"위로"` / `"아래로"` / `"필터"` / `"타이핑 검색어"` / `"Esc 로 NORMAL 모드"` 등) + mode-aware (NORMAL = navigation verbs, INSERT = typing + Esc reminder) + Library filter overlay 별 분기. 8 unit tests pin 모든 (pane, mode, filter) 조합 — exhaustive non-empty + Library Normal/filter, Search Normal/Insert, Ask Normal/Insert, Inspect Normal 별 verb fragment 존재 검증. spec status `in_progress` → `completed` — p9-fb-13 partial 의 deferred verb-form 항목이 닫힘.
|
||||
- **2026-05-03 P9 도그푸딩 후속 (p9-fb-13)** — TUI cheatsheet popup. `kebab-tui::cheatsheet::render_cheatsheet(f, area, app)` 신규 — 70%/60% centered modal, sections (Global / Library / Search / Ask / Inspect) + global toggle table + 현재 focused pane footer. `App.cheatsheet_visible: bool` 필드 + `pub fn cheatsheet_visible()` getter. run loop `cheatsheet_intercept(app, key)` 가 mode_intercept 보다 먼저 dispatch — `F1` 토글 (open/close), `Esc` 가 visible 일 때 닫기 (mode_intercept 를 우회해서 cheatsheet 닫기 가 mode flip 도 발동시키지 않도록), 그 외 키는 fall-through (popup 열린 채 navigation 가능). modifier-bearing F1 (Ctrl-F1 등) 은 무시. **HOTFIXES 기록**: spec 의 `?` trigger 가 Library 의 quick-Ask binding 과 충돌해서 `F1` 으로 rebind. spec 의 verb-form hint line 재구성은 별 후속 PR (기존 footer 가 동일 역할). spec status `planned` → `in_progress` (verb hint deferral 으로 partial). spec: `tasks/p9/p9-fb-13-tui-cheatsheet.md`.
|
||||
|
||||
## 다음 task 후보
|
||||
|
||||
|
||||
63
README.md
63
README.md
@@ -47,7 +47,7 @@ embedding 벡터를 청크 **내용 해시** 로 캐싱한다 (`derivation_cache
|
||||
|
||||
### 외부 계산 + 로컬 검색 워크플로
|
||||
|
||||
search/ask 는 원본 파일 없이 KB 산출물만으로 동작한다 (청크 본문이 SQLite 에 저장되고 문서 경로는 상대경로로 기록됨). 비싼 색인(임베딩·OCR)을 성능 좋은 머신에서 수행한 뒤(예: Apple Silicon 맥에서 candle Metal GPU), **두 산출물만** 다른 머신(예: NUMA 서버)으로 복사하면 그대로 검색·질문할 수 있다.
|
||||
search/ask 는 원본 파일 없이 KB 산출물만으로 동작한다 (청크 본문이 SQLite 에 저장되고 문서 경로는 상대경로로 기록됨). 비싼 색인(임베딩·OCR)을 성능 좋은 머신에서 수행한 뒤, **두 산출물만** 다른 머신으로 복사하면 그대로 검색·질문할 수 있다.
|
||||
|
||||
**무엇을 복사하나 — `[storage]` 에서 정의된 두 경로:**
|
||||
|
||||
@@ -64,7 +64,7 @@ rsync -a <src-data_dir>/kebab.sqlite user@server:<dst-data_dir>/
|
||||
rsync -a <src-data_dir>/lancedb/ user@server:<dst-data_dir>/lancedb/
|
||||
```
|
||||
|
||||
조건: **양쪽 동일 `kebab` 버전 + 동일 임베딩 모델/차원** (`[models.embedding].model`·`dimensions`). provider 는 달라도 됨 (예: 맥 `candle`/Metal ↔ 서버 `candle`/CPU 또는 `fastembed` — 같은 모델이면 벡터 호환). 복사는 반드시 ingest 가 돌지 않을 때.
|
||||
조건: **양쪽 동일 `kebab` 버전 + 동일 임베딩 모델/차원** (`[models.embedding].model`·`dimensions`). provider 는 달라도 됨 (예: `fastembed` ↔ `ollama` — 같은 모델이면 벡터 호환). 복사는 반드시 ingest 가 돌지 않을 때.
|
||||
|
||||
### 멀티미디어 색인
|
||||
|
||||
@@ -74,10 +74,6 @@ Markdown · PDF · 이미지(OCR + caption) · 소스코드(Rust/Python/TS/JS/Go
|
||||
|
||||
검색 결과를 근거로 LLM 답변을 생성하고 [#번호] 인용을 단다. 근거가 부족하면 답을 지어내지 않고 거절한다. compound 질문은 `--multi-hop` 으로 분해→synthesize. 답변의 groundedness 는 mDeBERTa XNLI 로 검증할 수 있다 (`[rag] nli_threshold`, default off).
|
||||
|
||||
### TUI
|
||||
|
||||
`kebab tui` 는 Ratatui 셸 — Library / Search / Ask / Inspect 패널을 vim-style 모드로 다룬다. 키 매핑은 앱 내 `F1` cheatsheet 가 권위 소스다.
|
||||
|
||||
## 명령
|
||||
|
||||
| 명령 | 동작 |
|
||||
@@ -94,7 +90,6 @@ Markdown · PDF · 이미지(OCR + caption) · 소스코드(Rust/Python/TS/JS/Go
|
||||
| `kebab eval run \| aggregate \| compare \| variants` | golden query 회귀 측정 + 변형 일관성 진단 |
|
||||
| `kebab schema [--json]` | introspection — wire schemas / capabilities / models / stats |
|
||||
| `kebab doctor` | 설정 / 모델 / DB 헬스 체크 |
|
||||
| `kebab tui` | Ratatui 셸 (Library / Search / Ask / Inspect) |
|
||||
| `kebab mcp` | MCP stdio server (`search` / `bulk_search` / `ask` / `fetch` / `schema` / `doctor` / `ingest_file` / `ingest_stdin`) |
|
||||
| `kebab reset [--all \| --data-only \| --vector-only \| --config-only \| --orphans-only] [--yes]` | XDG 데이터 wipe (**irreversible**) |
|
||||
|
||||
@@ -123,23 +118,16 @@ root = "~/KnowledgeBase" # 색인할 폴더. 절대 / tilde / env / 상대 경
|
||||
# trust_level = "secondary" # 낮은 신뢰 출처 — `--trust-min primary` 로 배제 가능.
|
||||
|
||||
[models.embedding]
|
||||
provider = "fastembed" # "fastembed"(기본, onnxruntime) / "candle"(순수 Rust)
|
||||
# / "ollama"(원격 HTTP) / "none"(lexical-only).
|
||||
# candle 는 같은 모델·같은 벡터를 순수 Rust 로 돌려
|
||||
# NUMA 서버의 onnxruntime 48-스레드 double-free 를 피하는
|
||||
# opt-in 백엔드 (e5 는 재색인 불필요).
|
||||
provider = "fastembed" # "fastembed"(기본, onnxruntime) / "ollama"(원격 HTTP)
|
||||
# / "none"(lexical-only).
|
||||
model = "multilingual-e5-large" # 다국어 sentence embedding (1024-dim).
|
||||
# 첫 ingest 시 ONNX (~1.3GB) 자동 다운로드.
|
||||
# candle provider 는 safetensors (~2GB) 다운로드.
|
||||
# candle/ollama 는 "snowflake-arctic-embed-l-v2.0"
|
||||
# ollama 는 "snowflake-arctic-embed-l-v2.0"
|
||||
# (설명형 query 의 recall 보강) 도 지원 — 아래 참고.
|
||||
dimensions = 1024 # config 와 LanceDB stored dim 불일치 시 검색 0건.
|
||||
num_threads = 0 # candle 전용 CPU 스레드 캡 (0=auto=#cores).
|
||||
# env KEBAB_EMBED_THREADS 가 우선. NUMA 노드 바인딩은
|
||||
# numactl 과 조합. fastembed provider 는 무시.
|
||||
# endpoint = "http://127.0.0.1:11434" # provider="ollama" 전용 HTTP endpoint.
|
||||
# 생략 시 [models.llm].endpoint 로 폴백.
|
||||
# fastembed/candle provider 는 무시.
|
||||
# fastembed provider 는 무시.
|
||||
```
|
||||
|
||||
**arctic-embed-l-v2.0 (설명형 query recall 보강)**: 기본 e5-large 대신
|
||||
@@ -147,13 +135,7 @@ Snowflake `arctic-embed-l-v2.0` 임베더를 쓸 수 있다 (1024-dim, opt-in).
|
||||
설명형/약어/영문 용어 query 의 recall@10 이 e5 대비 향상됐다. 두 경로:
|
||||
|
||||
```toml
|
||||
# (A) candle 백엔드 — 순수 Rust, in-process (NUMA 안전, Metal GPU 가능):
|
||||
[models.embedding]
|
||||
provider = "candle"
|
||||
model = "snowflake-arctic-embed-l-v2.0" # CLS pooling, query 에 "query: " 접두어
|
||||
# (문서는 무접두어). safetensors ~2GB 다운로드.
|
||||
|
||||
# (B) ollama 백엔드 — 원격/로컬 Ollama 데몬에 위임 (POST /api/embed):
|
||||
# ollama 백엔드 — 원격/로컬 Ollama 데몬에 위임 (POST /api/embed):
|
||||
[models.embedding]
|
||||
provider = "ollama"
|
||||
model = "snowflake-arctic-embed2" # Ollama 모델 태그 (ollama pull 필요)
|
||||
@@ -164,21 +146,6 @@ endpoint = "http://127.0.0.1:11434" # 생략 시 [models.llm].endpoint
|
||||
> 벡터도 다름). 기존 e5 KB 와 혼용 불가 — 전환 시 **재색인** 필요 (`kebab reset`
|
||||
> 후 재 ingest). 기본값은 e5 라 기존 사용자는 영향 없음.
|
||||
|
||||
**Apple Silicon GPU 가속 (candle / macOS)**: M-시리즈 맥에서 candle 임베딩을
|
||||
GPU(Metal)로 돌리면 CPU 대비 대용량 ingest 가 크게 빨라진다. 빌드 또는 설치 시
|
||||
`embed_metal` feature 를 켠다:
|
||||
|
||||
```bash
|
||||
# 빌드만:
|
||||
cargo build --release --features embed_metal
|
||||
# 전역 설치 (~/.cargo/bin/kebab):
|
||||
cargo install --path crates/kebab-cli --features embed_metal --locked
|
||||
```
|
||||
|
||||
벡터는 CPU candle 과 동일 모델이라 호환되므로, 맥에서 GPU 로 색인한
|
||||
`kebab.sqlite` + `lancedb/` 를 그대로 Linux 서버(CPU candle)로 복사해 질의할 수
|
||||
있다. 색인 로그에 `candle device = Metal (GPU)` 가 보이면 GPU 사용 중. metal
|
||||
feature 는 macOS 전용 (Linux/서버는 기본 CPU 빌드).
|
||||
|
||||
```toml
|
||||
|
||||
@@ -191,7 +158,7 @@ model = "gemma4:e4b"
|
||||
stale_threshold_days = 30 # search hit / citation 의 stale 플래그 기준 (0 = off).
|
||||
|
||||
[rag]
|
||||
prompt_template_version = "rag-v4" # 각 근거의 source/trust 라벨로 low-trust 출처 discount + 답변 언어 = 질문 언어. rag-v1/v2/v3 는 legacy.
|
||||
prompt_template_version = "rag-v4" # 각 근거의 source/trust 라벨로 low-trust 출처 discount + 답변 언어 = 질문 언어. rag-v3 는 legacy.
|
||||
nli_threshold = 0.0 # >0 (예: 0.5) 면 mDeBERTa XNLI groundedness 검증.
|
||||
```
|
||||
|
||||
@@ -199,11 +166,12 @@ nli_threshold = 0.0 # >0 (예: 0.5) 면 mDeBERTa XNLI groundedn
|
||||
- **`[ingest.chunking]`** — 청크 크기·오버랩·heading 존중. `chunker_version` 기본 `"md-heading-v2"` (v0.30.0). **`max_chunk_tokens`** (default 4000, byte/3 토큰) — 이 값을 넘는 청크는 줄(→UTF-8 char) 경계로 분할해 각 조각이 예산 이하가 되게 한다. 거대 list/code/log 덤프가 한 청크로 임베더 컨텍스트를 초과하던 문제를 막는다(미분할 청크는 v0.29.0 `md-heading-v1` 과 출력 동일). 이 값을 바꾸면 markdown 자산이 자동 재청크된다.
|
||||
- **파생물 캐시** — embedding 결과를 내용 해시로 자동 캐싱한다 (위 「핵심 기능」 참고). 설정 항목 없음.
|
||||
- **`[ingest.code]`** — code ingest 의 skip 정책 (`skip_generated_header`, `max_file_bytes`, `extra_skip_globs`). `.gitignore` 자동 honor, `.kebabignore` 는 추가 layer.
|
||||
- **`[ingest.image.ocr]`** — 이미지 OCR (default off / opt-in). `engine` 으로 백엔드 선택: `"ollama-vision"` (default, 원격 vision LM) 또는 `"paddle-onnx"` (PP-OCRv5 ONNX 를 in-process 로 실행, Python 런타임 불필요, 큰 페이지 CPU <4초, 오프라인). `paddle-onnx` 는 워크스페이스에 번들된 모델을 쓰며 `det_model`/`rec_model`/`dict` 로 경로 override, `score_thresh`(0.3)/`unclip_ratio`(1.5)/`max_boxes`(1000) 로 검출 튜닝 가능 (`KEBAB_IMAGE_OCR_*` env 동일 지원 — env 이름은 v3 에서도 불변). engine 또는 모델을 바꾸면 영향 이미지가 자동 재색인된다.
|
||||
- **`[ingest.pdf.ocr]`** — scanned PDF 의 page-단위 OCR (default off / opt-in, page 당 ~수십 초 cost). `engine` 은 `[ingest.image.ocr]` 과 동일하게 `"ollama-vision"`/`"paddle-onnx"` 선택. v3 에서 paddle 모델 경로 키(`det_model`/`rec_model`/`dict`/`score_thresh`/`unclip_ratio`/`max_boxes`)를 PDF 자체적으로 가질 수 있다(`KEBAB_PDF_OCR_*` env 동일). 활성화 후 옛 색인분은 `kebab ingest --force-reingest` 로 재처리.
|
||||
- **`--config <path>`** — 임시 워크스페이스 / 격리 테스트용 (CLI · TUI 모두 honor).
|
||||
- **`[ingest.ocr]`** (config schema v5) — image/pdf OCR 가 공유하는 **엔진** 설정의 단일 출처 (`engine`/`model`/`endpoint`/`languages`/`max_pixels`/`request_timeout_secs` + paddle 모델 경로·튜닝 키). 여기에 한 번 적어 두면 image·pdf 양쪽에 적용되고, 각 미디어 블록(`[ingest.image.ocr]`/`[ingest.pdf.ocr]`)이 자기 키로 override 한다 (우선순위: 미디어 블록 > `[ingest.ocr]` > 내장 기본값). 옛 v4 `config.toml` 의 image/pdf 에 중복돼 있던 OCR 엔진 키는 로드 시 자동으로 이 블록으로 통합된다 (effective 값 불변, 자동 재색인 없음). env override 도 `KEBAB_OCR_*` 하나로 통합 (양쪽 미디어에 적용).
|
||||
- **`[ingest.image.ocr]`** — 이미지 OCR. on/off 토글(`enabled`, default off / opt-in)은 미디어별이며, 엔진 설정은 `[ingest.ocr]` 에서 상속하되 이 블록에서 override 할 수 있다. `engine` 으로 백엔드 선택: `"ollama-vision"` (default, 원격 vision LM) 또는 `"paddle-onnx"` (PP-OCRv5 ONNX 를 in-process 로 실행, Python 런타임 불필요, 큰 페이지 CPU <4초, 오프라인). `paddle-onnx` 는 워크스페이스에 번들된 모델을 쓰며 `det_model`/`rec_model`/`dict` 로 경로 override, `score_thresh`(0.3)/`unclip_ratio`(1.5)/`max_boxes`(1000) 로 검출 튜닝 가능. engine 또는 모델을 바꾸면 영향 이미지가 자동 재색인된다.
|
||||
- **`[ingest.pdf.ocr]`** — scanned PDF 의 page-단위 OCR (default off / opt-in, page 당 ~수십 초 cost). on/off 토글(`enabled`/`always_on`)과 PDF 고유 키(`valid_ratio_threshold`/`min_char_count`/`lang_hint`)는 미디어별이고, 엔진 설정은 `[ingest.ocr]` 에서 상속하되 이 블록에서 override 한다(PDF 기본 모델은 `qwen2.5vl:3b`, 이미지의 `gemma4:e4b` 와 다름 — 미디어별 기본값 보존). 활성화 후 옛 색인분은 `kebab ingest --force-reingest` 로 재처리.
|
||||
- **`--config <path>`** — 임시 워크스페이스 / 격리 테스트용 (CLI honor).
|
||||
- **`kebab config migrate`** — 새 버전에서 추가된 config 섹션을 기존 `config.toml` 에 설명 주석과 함께 채워 넣는다 (사용자가 손본 값·주석·순서는 보존, 멱등, 변경 시 자동 `.bak` 백업). `--dry-run` 으로 변경 미리보기. `kebab doctor` 가 갱신 필요 시 안내한다. `kebab init` 으로 새로 생성되는 config.toml 도 섹션별 주석을 포함한다.
|
||||
- **`KEBAB_*` env** — 일부 키 override (`KEBAB_RAG_SCORE_GATE`, `KEBAB_EVAL_GOLDEN` 등).
|
||||
- **`KEBAB_*` env** — 런타임 override용 ~22개 키만 노출. 엔드포인트(`KEBAB_MODELS_LLM_ENDPOINT`, `KEBAB_MODELS_EMBEDDING_ENDPOINT`, `KEBAB_OCR_ENDPOINT`), 모델명/프로바이더(`KEBAB_MODELS_LLM_MODEL`, `KEBAB_MODELS_EMBEDDING_MODEL`, `KEBAB_MODELS_EMBEDDING_PROVIDER`, `KEBAB_MODELS_LLM_PROVIDER`, `KEBAB_MODELS_NLI_MODEL`), 경로(`KEBAB_WORKSPACE_ROOT`, `KEBAB_STORAGE_DATA_DIR`), 병렬도(`KEBAB_INDEXING_MAX_PARALLEL_EXTRACTORS`, `KEBAB_INDEXING_MAX_PARALLEL_EMBEDDINGS`), 청킹(`KEBAB_CHUNKING_TARGET_TOKENS`, `KEBAB_CHUNKING_OVERLAP_TOKENS`), OCR 토글/엔진/언어(`KEBAB_IMAGE_OCR_ENABLED`, `KEBAB_PDF_OCR_ENABLED`, `KEBAB_OCR_ENGINE`, `KEBAB_OCR_MODEL`, `KEBAB_OCR_LANGUAGES`), 기타(`KEBAB_IMAGE_CAPTION_ENABLED`, `KEBAB_SEARCH_DEFAULT_K`, `KEBAB_RAG_PROMPT_TEMPLATE_VERSION`). 나머지 세부 튜닝 키(score_gate, rrf_k, temperature 등)는 `config.toml` 전용. 특수: `KEBAB_READONLY=1`(write-path 비활성), `KEBAB_PROGRESS=plain`(non-TTY 진행 출력), `KEBAB_EVAL_GOLDEN`(eval golden set 경로).
|
||||
- **XDG layout**: `~/.config/kebab/`, `~/.local/share/kebab/`, `~/.cache/kebab/`, `~/.local/state/kebab/`.
|
||||
|
||||
## 아키텍처
|
||||
@@ -214,7 +182,6 @@ flowchart TB
|
||||
|
||||
subgraph UI["UI binary"]
|
||||
cli["kebab CLI"]
|
||||
tui["kebab TUI"]
|
||||
end
|
||||
|
||||
subgraph App["Facade"]
|
||||
@@ -241,9 +208,7 @@ flowchart TB
|
||||
end
|
||||
|
||||
user --> cli
|
||||
user --> tui
|
||||
cli --> app
|
||||
tui --> app
|
||||
|
||||
app --> parse
|
||||
app --> chunker
|
||||
@@ -267,7 +232,7 @@ flowchart TB
|
||||
|
||||
v0.21.0 기준 핵심 설계:
|
||||
|
||||
- **crate facade** — `kebab-app` 가 유일한 facade다. UI binary (`kebab-cli` / `kebab-tui`) 는 store / parse / search / llm / rag 를 직접 참조하지 않는다 (frozen 설계 §8). 각 user-facing 엔트리는 `*_with_config(cfg, …)` 동반 함수로 explicit config 를 thread 한다.
|
||||
- **crate facade** — `kebab-app` 가 유일한 facade다. UI binary (`kebab-cli`) 는 store / parse / search / llm / rag 를 직접 참조하지 않는다 (frozen 설계 §8). 각 user-facing 엔트리는 `*_with_config(cfg, …)` 동반 함수로 explicit config 를 thread 한다.
|
||||
- **chunk_id 는 위치 기반** — chunk 의 정체성은 문서 내 위치(ordinal + span)다. 반면 파생물 캐시 키는 **내용 해시**라, 내용이 같으면 위치·문서가 달라도 동일 캐시를 재사용한다.
|
||||
- **wire schema v1** — 모든 `--json` 출력은 `schema_version` 을 담는 frozen contract다. 깨는 변경은 `*.v2` major bump을 요구한다.
|
||||
- **versioning cascade** — `parser_version` / `chunker_version` / `embedding_version` / `prompt_template_version` / `index_version` 변경은 downstream record(청크·임베딩·캐시·eval)를 무효화한다.
|
||||
|
||||
@@ -18,7 +18,6 @@ kebab-store-vector = { path = "../kebab-store-vector" }
|
||||
kebab-search = { path = "../kebab-search" }
|
||||
kebab-embed = { path = "../kebab-embed" }
|
||||
kebab-embed-local = { path = "../kebab-embed-local" }
|
||||
kebab-embed-candle = { path = "../kebab-embed-candle" }
|
||||
kebab-embed-ollama = { path = "../kebab-embed-ollama" }
|
||||
kebab-llm = { path = "../kebab-llm" }
|
||||
kebab-llm-local = { path = "../kebab-llm-local" }
|
||||
@@ -57,12 +56,6 @@ tracing-subscriber = { version = "0.3", features = ["env-filter", "fmt", "json
|
||||
tracing-appender = "0.2"
|
||||
toml = "0.8"
|
||||
dirs = "5"
|
||||
# p9-fb-19: in-process LRU cache for `App::search`. Capacity from
|
||||
# `config.search.cache_capacity` (default 256, ~1.3 MB cap).
|
||||
lru = { workspace = true }
|
||||
# p9-fb-19: NFKC-normalize cache-key queries so `"Foo"` / `"FOO"` /
|
||||
# `" foo "` collapse to one entry. Same crate kebab-normalize +
|
||||
# kebab-core already use, no version drift.
|
||||
unicode-normalization = "0.1"
|
||||
# p9-fb-31: GitignoreBuilder for .kebabignore matching in ingest_file_with_config.
|
||||
# Same version as kebab-source-fs (0.4) to avoid duplicate dep versions.
|
||||
@@ -101,8 +94,6 @@ reqwest = { version = "0.12", default-features = false, features = ["blocki
|
||||
# disable path 없음; 이 feature 는 spec §6.3 명시를 honor 하는 role 만.
|
||||
default = ["fts_korean_morphological"]
|
||||
fts_korean_morphological = []
|
||||
# opt-in (macOS): candle embedder runs on the Apple Silicon GPU. See kebab-embed-candle.
|
||||
embed_metal = ["kebab-embed-candle/metal"]
|
||||
|
||||
[lints]
|
||||
workspace = true
|
||||
|
||||
@@ -33,18 +33,15 @@
|
||||
//! in that mode [`App::embedder`] returns `None` and callers must fall
|
||||
//! back to lexical-only search.
|
||||
|
||||
use std::num::NonZeroUsize;
|
||||
use std::sync::{Arc, Mutex, OnceLock};
|
||||
use std::sync::{Arc, OnceLock};
|
||||
|
||||
use anyhow::{Context, Result, anyhow};
|
||||
use lru::LruCache;
|
||||
|
||||
use kebab_core::{
|
||||
Answer, DocumentStore, Embedder, ExtractContext, Extractor, IndexVersion, LanguageModel,
|
||||
MediaType, Retriever, SearchHit, SearchMode, SearchOpts, SearchQuery, VectorStore,
|
||||
};
|
||||
use kebab_embed_candle::CandleEmbedder;
|
||||
use kebab_embed_local::FastembedEmbedder;
|
||||
use kebab_embed_local::{FASTEMBED_CACHE_SUBDIR, FastembedEmbedder};
|
||||
use kebab_embed_ollama::OllamaEmbedder;
|
||||
use kebab_llm_local::OllamaLanguageModel;
|
||||
use kebab_parse_code::{
|
||||
@@ -52,6 +49,7 @@ use kebab_parse_code::{
|
||||
KotlinAstExtractor, PythonAstExtractor, RustAstExtractor, TypescriptAstExtractor,
|
||||
};
|
||||
use kebab_parse_image::ImageExtractor;
|
||||
use kebab_parse_md::MarkdownExtractor;
|
||||
use kebab_parse_pdf::PdfTextExtractor;
|
||||
use kebab_rag::{AskOpts, RagPipeline};
|
||||
use kebab_search::{HybridRetriever, LexicalRetriever, VectorRetriever};
|
||||
@@ -102,9 +100,9 @@ pub struct App {
|
||||
pub(crate) sqlite: Arc<SqliteStore>,
|
||||
/// post-v0.18.0 extractor-dispatch-unification: polymorphic Extractor
|
||||
/// registry. App init 시 1회 등록되어 `extract_for(...)` 가 lookup
|
||||
/// 한다. 현재 11 entry (ImageExtractor + PdfTextExtractor + 9 AST).
|
||||
/// MarkdownExtractor 는 별 PR 에서 추가 — markdown ingest path 는
|
||||
/// 본 PR 에서 free-function 그대로 유지.
|
||||
/// 한다. 현재 12 entry (MarkdownExtractor + ImageExtractor +
|
||||
/// PdfTextExtractor + 9 AST). MarkdownExtractor 가 마지막으로 합류해
|
||||
/// 모든 media 가 `extract_for` 경유로 통일됨 (extract-stage 대칭화).
|
||||
pub(crate) extractors: Vec<Box<dyn Extractor + Send + Sync>>,
|
||||
/// Memoized embedder — built lazily on first `embedder()` call when
|
||||
/// embeddings are enabled. `OnceLock` keeps the struct `Sync` and
|
||||
@@ -118,16 +116,10 @@ pub struct App {
|
||||
/// client per query (cheap, but still measurable on a 50-query
|
||||
/// suite).
|
||||
llm: OnceLock<Arc<dyn LanguageModel>>,
|
||||
/// p9-fb-19: in-process LRU search-result cache. Capacity comes
|
||||
/// from `config.search.cache_capacity` (default 256, ~1.3 MB
|
||||
/// cap). `None` when capacity is 0 (cache disabled). The
|
||||
/// `corpus_revision` snapshot embedded in `SearchCacheKey`
|
||||
/// invalidates every entry the moment a new ingest commit lands.
|
||||
search_cache: Option<Mutex<LruCache<SearchCacheKey, Vec<SearchHit>>>>,
|
||||
/// p9-fb-41 PR-9c-2: NLI verifier built eagerly at
|
||||
/// `open_with_config` time when `config.rag.nli_threshold > 0`,
|
||||
/// consumed by `RagPipeline::with_verifier` on every `ask` /
|
||||
/// `ask_with_session` call. `None` when the gate is disabled
|
||||
/// consumed by `RagPipeline::with_verifier` on every `ask` call.
|
||||
/// `None` when the gate is disabled
|
||||
/// (default, threshold = 0) — multi-hop skips step 8.5 entirely
|
||||
/// and single-pass never touches the verifier.
|
||||
///
|
||||
@@ -137,46 +129,6 @@ pub struct App {
|
||||
pipeline_verifier: Option<Arc<dyn kebab_nli::NliVerifier>>,
|
||||
}
|
||||
|
||||
/// p9-fb-19: cache key for `App::search`. Includes every field that
|
||||
/// could change the result set:
|
||||
/// - normalized query (NFKC + trim + lowercase)
|
||||
/// - mode + k + snippet_chars (caller knobs)
|
||||
/// - embedding_version + chunker_version (model identity)
|
||||
/// - corpus_revision (monotonic counter that ingest bumps)
|
||||
///
|
||||
/// Lexical mode has no embedding identity → empty string in that
|
||||
/// slot, harmless because the rest of the key still distinguishes
|
||||
/// queries.
|
||||
///
|
||||
/// **Naming note**: spec p9-fb-19 calls the invalidation counter
|
||||
/// `index_version`, but the impl renames it to `corpus_revision` to
|
||||
/// avoid confusion with the pre-existing `IndexVersion` newtype
|
||||
/// (design §9 — embedding-index identity label, a completely
|
||||
/// different concept). The `corpus_revision` row in the §9
|
||||
/// versioning table documents the new dimension; HOTFIXES entry
|
||||
/// tracks the rename.
|
||||
#[derive(Clone, Debug, Eq, Hash, PartialEq)]
|
||||
pub(crate) struct SearchCacheKey {
|
||||
pub query_norm: String,
|
||||
pub mode: SearchMode,
|
||||
pub k: u32,
|
||||
pub snippet_chars: u32,
|
||||
pub embedding_version: String,
|
||||
pub chunker_version: String,
|
||||
pub corpus_revision: u64,
|
||||
}
|
||||
|
||||
impl SearchCacheKey {
|
||||
/// Normalize `query.text` per spec p9-fb-19: NFKC + trim +
|
||||
/// lowercase. Means `"Foo"` / `"FOO"` / `" foo "` collapse to a
|
||||
/// single cache entry — redundant work avoided when the user's
|
||||
/// input differs only in shape.
|
||||
pub fn normalize_query(text: &str) -> String {
|
||||
use unicode_normalization::UnicodeNormalization;
|
||||
text.trim().nfkc().collect::<String>().to_lowercase()
|
||||
}
|
||||
}
|
||||
|
||||
impl App {
|
||||
/// Open the SQLite store and run migrations. Does NOT load the
|
||||
/// embedder or vector store — those are lazy via
|
||||
@@ -187,7 +139,7 @@ impl App {
|
||||
/// internally drives a `tokio::Runtime::block_on`, which panics if
|
||||
/// invoked from inside another tokio runtime.
|
||||
pub fn open_with_config(config: kebab_config::Config) -> Result<Self> {
|
||||
let sqlite = SqliteStore::open(&config).context("kb-app: open SqliteStore")?;
|
||||
let sqlite = SqliteStore::open(&config.storage).context("kb-app: open SqliteStore")?;
|
||||
sqlite
|
||||
.run_migrations()
|
||||
.context("kb-app: run SqliteStore migrations")?;
|
||||
@@ -219,17 +171,16 @@ impl App {
|
||||
"korean tokenizer backfill complete: {backfill_count} chunks updated"
|
||||
);
|
||||
}
|
||||
// p9-fb-19: build the LRU cache from config. Capacity 0 →
|
||||
// `None` (cache disabled — every search hits the retrievers).
|
||||
let search_cache = NonZeroUsize::new(config.search.cache_capacity)
|
||||
.map(|cap| Mutex::new(LruCache::new(cap)));
|
||||
// post-v0.18.0 extractor-dispatch-unification: build the 11-entry
|
||||
// post-v0.18.0 extractor-dispatch-unification: build the 12-entry
|
||||
// Extractor registry. All entries are state-less unit structs with
|
||||
// zero-cost `new()`, so init cost is effectively 0 and side effects
|
||||
// are 0 — `pipeline_verifier` fallible `?` below may bail but the
|
||||
// already-constructed `extractors` Vec drops without cost. Markdown
|
||||
// is NOT registered (see field doc).
|
||||
// already-constructed `extractors` Vec drops without cost.
|
||||
// MarkdownExtractor is registered first so markdown ingest flows
|
||||
// through `extract_for` like every other media (extract-stage
|
||||
// symmetry — previously the only free-function arm).
|
||||
let extractors: Vec<Box<dyn Extractor + Send + Sync>> = vec![
|
||||
Box::new(MarkdownExtractor::new()),
|
||||
Box::new(ImageExtractor::new()),
|
||||
Box::new(PdfTextExtractor::new()),
|
||||
Box::new(RustAstExtractor::new()),
|
||||
@@ -264,7 +215,6 @@ impl App {
|
||||
embedder: OnceLock::new(),
|
||||
vector: OnceLock::new(),
|
||||
llm: OnceLock::new(),
|
||||
search_cache,
|
||||
pipeline_verifier,
|
||||
})
|
||||
}
|
||||
@@ -294,73 +244,13 @@ impl App {
|
||||
}
|
||||
|
||||
/// Run a [`SearchQuery`] through the configured retriever stack and
|
||||
/// return the top-k hits. p9-fb-19: result is served from the
|
||||
/// in-process LRU cache when the same `(query_norm, mode, k,
|
||||
/// snippet_chars, embedding_version, chunker_version,
|
||||
/// corpus_revision)` tuple was seen before; cache miss falls
|
||||
/// through to [`Self::search_uncached`].
|
||||
/// return the top-k hits.
|
||||
///
|
||||
/// Reuses any previously-built embedder / vector store on this `App`
|
||||
/// — long-lived callers (kb-eval, future TUI) get amortized cost
|
||||
/// across calls.
|
||||
pub fn search(&self, query: SearchQuery) -> Result<Vec<SearchHit>> {
|
||||
let Some(cache) = self.search_cache.as_ref() else {
|
||||
// Cache disabled (capacity = 0) — straight-line.
|
||||
return self.search_uncached(query);
|
||||
};
|
||||
// Build the cache key. embedding_version is empty for lexical
|
||||
// mode (no embedder identity); for vector/hybrid we need the
|
||||
// embedder built (which forces the cold-start cost), but
|
||||
// that's the cost the cache exists to amortize across
|
||||
// *subsequent* identical queries.
|
||||
let key = self.build_cache_key(&query)?;
|
||||
// Lock the cache long enough to lookup; clone the hit out so
|
||||
// we can drop the lock before returning. Mutex poison
|
||||
// recovery: `into_inner()` of a poison error returns the
|
||||
// (still-valid) underlying guard so we can keep using the
|
||||
// cache after a panic in another thread. Log once so the
|
||||
// poison itself is visible — the cache is still functional
|
||||
// but a panic in a previous search is worth knowing about.
|
||||
let mut guard = cache.lock().unwrap_or_else(|e| {
|
||||
tracing::warn!(
|
||||
target: "kebab-app",
|
||||
"search_cache mutex was poisoned; recovering and continuing — \
|
||||
a previous search-thread panic preceded this call"
|
||||
);
|
||||
e.into_inner()
|
||||
});
|
||||
if let Some(hits) = guard.get(&key) {
|
||||
tracing::debug!(
|
||||
target: "kebab-app",
|
||||
cache = "hit",
|
||||
corpus_revision = key.corpus_revision,
|
||||
"search served from LRU cache"
|
||||
);
|
||||
// p9-fb-32: re-stamp staleness on every cache hit. The cache
|
||||
// entry was stamped at insert time against an older `now`
|
||||
// and an older threshold; if either has shifted (config
|
||||
// reload, time passing) the cached `stale: false` may now
|
||||
// be wrong. Re-stamping is cheap (per-hit comparison) and
|
||||
// avoids invalidating the cache on threshold changes.
|
||||
let mut hits = hits.clone();
|
||||
drop(guard);
|
||||
let now = time::OffsetDateTime::now_utc();
|
||||
crate::staleness::mark_stale_in_place(
|
||||
&mut hits,
|
||||
now,
|
||||
self.config.search.stale_threshold_days,
|
||||
);
|
||||
return Ok(hits);
|
||||
}
|
||||
// Drop the lock before the (potentially slow) retriever call
|
||||
// so other in-flight searches can use the cache concurrently.
|
||||
drop(guard);
|
||||
let hits = self.search_uncached(query)?;
|
||||
let mut guard = cache
|
||||
.lock()
|
||||
.unwrap_or_else(std::sync::PoisonError::into_inner);
|
||||
guard.put(key, hits.clone());
|
||||
Ok(hits)
|
||||
self.search_uncached(query)
|
||||
}
|
||||
|
||||
/// p9-fb-19: bypass the LRU cache and run the search directly.
|
||||
@@ -407,7 +297,7 @@ impl App {
|
||||
vec_iv,
|
||||
self.config.search.snippet_chars,
|
||||
)) as Arc<dyn Retriever>;
|
||||
let hybrid = HybridRetriever::new(&self.config, lex, vec_retr);
|
||||
let hybrid = HybridRetriever::new(&self.config.search, lex, vec_retr);
|
||||
hybrid.search(&query)?
|
||||
}
|
||||
};
|
||||
@@ -505,7 +395,7 @@ impl App {
|
||||
self.config.search.snippet_chars,
|
||||
)) as Arc<dyn Retriever>
|
||||
};
|
||||
let hybrid = HybridRetriever::new(&self.config, lex, vec_retr);
|
||||
let hybrid = HybridRetriever::new(&self.config.search, lex, vec_retr);
|
||||
let (mut traced_hits, trace) = hybrid.search_with_trace(&fetch_query)?;
|
||||
|
||||
// Stamp staleness — same as search_uncached.
|
||||
@@ -639,8 +529,8 @@ impl App {
|
||||
pipeline.ask(query, opts)
|
||||
}
|
||||
|
||||
/// p9-fb-41 PR-9c-2: shared pipeline builder used by [`Self::ask`]
|
||||
/// and [`Self::ask_with_session`]. Attaches the App-built NLI
|
||||
/// p9-fb-41 PR-9c-2: shared pipeline builder used by [`Self::ask`].
|
||||
/// Attaches the App-built NLI
|
||||
/// verifier (when `cfg.rag.nli_threshold > 0`) via
|
||||
/// `RagPipeline::with_verifier`, keeping the construction site in
|
||||
/// a single place so the two call paths can't drift.
|
||||
@@ -649,15 +539,21 @@ impl App {
|
||||
retriever: Arc<dyn Retriever>,
|
||||
llm: Arc<dyn LanguageModel>,
|
||||
) -> RagPipeline {
|
||||
let pipeline = RagPipeline::new(self.config.clone(), retriever, llm, self.sqlite.clone());
|
||||
let pipeline = RagPipeline::new(
|
||||
self.config.rag.clone(),
|
||||
self.config.models.clone(),
|
||||
self.config.search.clone(),
|
||||
retriever,
|
||||
llm,
|
||||
self.sqlite.clone(),
|
||||
);
|
||||
match &self.pipeline_verifier {
|
||||
Some(v) => pipeline.with_verifier(v.clone()),
|
||||
None => pipeline,
|
||||
}
|
||||
}
|
||||
|
||||
/// p9-fb-18: shared retriever-stack builder used by [`Self::ask`]
|
||||
/// and [`Self::ask_with_session`]. Lexical mode uses the FTS5
|
||||
/// Shared retriever-stack builder used by [`Self::ask`]. Lexical mode uses the FTS5
|
||||
/// retriever directly; vector / hybrid require embeddings (and
|
||||
/// surface the same "switch to --mode lexical" error from
|
||||
/// [`Self::require_embeddings`] when disabled).
|
||||
@@ -698,124 +594,11 @@ impl App {
|
||||
vec_iv,
|
||||
self.config.search.snippet_chars,
|
||||
)) as Arc<dyn Retriever>;
|
||||
Arc::new(HybridRetriever::new(&self.config, lex, vec_retr))
|
||||
Arc::new(HybridRetriever::new(&self.config.search, lex, vec_retr))
|
||||
}
|
||||
})
|
||||
}
|
||||
|
||||
/// p9-fb-18: ask under a persistent chat session. Loads the
|
||||
/// session's prior turns (if any), runs the query through
|
||||
/// `RagPipeline::ask_with_history`, then appends the new turn
|
||||
/// + (auto-)creates the session row on first use.
|
||||
///
|
||||
/// `session_id` is caller-supplied. If the session doesn't
|
||||
/// exist yet, a new `chat_sessions` row is created with title
|
||||
/// derived from the first question (≤40 chars, trimmed and
|
||||
/// NFC-normalized). Subsequent calls with the same
|
||||
/// `session_id` extend the conversation.
|
||||
///
|
||||
/// The returned `Answer` carries `conversation_id = Some(
|
||||
/// session_id)` and `turn_index = Some(n)` per p9-fb-15. The
|
||||
/// new `chat_turns` row is committed before this method
|
||||
/// returns; on persistence error, the answer is still returned
|
||||
/// (don't lose the user's compute) but the error is logged so
|
||||
/// the operator notices.
|
||||
pub fn ask_with_session(&self, session_id: &str, query: &str, opts: AskOpts) -> Result<Answer> {
|
||||
use kebab_core::traits::{ChatSessionRepo, ChatSessionRow, ChatTurnRow};
|
||||
use std::time::{SystemTime, UNIX_EPOCH};
|
||||
|
||||
// Load (or create) the session header.
|
||||
let now_unix = SystemTime::now()
|
||||
.duration_since(UNIX_EPOCH)
|
||||
.map_or(0, |d| d.as_secs() as i64);
|
||||
let existing = self.sqlite.get_session(session_id)?;
|
||||
let prior_turns = match &existing {
|
||||
Some(_) => self.sqlite.list_turns(session_id)?,
|
||||
None => Vec::new(),
|
||||
};
|
||||
let next_index = u32::try_from(prior_turns.len()).unwrap_or(u32::MAX);
|
||||
|
||||
// Build history Vec<Turn> from the persisted rows. Citations
|
||||
// are decoded best-effort — a corrupted citations_json
|
||||
// becomes an empty Vec rather than a panic (history is
|
||||
// advisory, not authoritative).
|
||||
let history: Vec<kebab_core::Turn> = prior_turns
|
||||
.iter()
|
||||
.map(|row| kebab_core::Turn {
|
||||
question: row.question.clone(),
|
||||
answer: row.answer.clone(),
|
||||
citations: serde_json::from_str(&row.citations_json).unwrap_or_default(),
|
||||
created_at: time::OffsetDateTime::from_unix_timestamp(row.created_at)
|
||||
.unwrap_or(time::OffsetDateTime::UNIX_EPOCH),
|
||||
})
|
||||
.collect();
|
||||
|
||||
// p9-fb-18 R1: shared retriever builder removes the prior
|
||||
// copy of `ask`'s 35-line stack — see [`Self::build_retriever`].
|
||||
// p9-fb-41 PR-9c-2: shared `build_pipeline` attaches the NLI
|
||||
// verifier when the gate is enabled.
|
||||
let retriever = self.build_retriever(opts.mode)?;
|
||||
let llm = self.llm()?;
|
||||
let pipeline = self.build_pipeline(retriever, llm);
|
||||
let answer =
|
||||
pipeline.ask_with_history(query, history, session_id.to_string(), next_index, opts)?;
|
||||
|
||||
// Auto-create the session header on first use. Title from
|
||||
// the first question (≤40 chars after trim).
|
||||
if existing.is_none() {
|
||||
let title = first_question_title(query);
|
||||
let session_row = ChatSessionRow {
|
||||
session_id: session_id.to_string(),
|
||||
created_at: now_unix,
|
||||
updated_at: now_unix,
|
||||
title: Some(title),
|
||||
config_snapshot_json: serde_json::json!({
|
||||
"prompt_template_version": self.config.rag.prompt_template_version,
|
||||
"llm.model": self.config.models.llm.model,
|
||||
"max_context_tokens": self.config.rag.max_context_tokens,
|
||||
})
|
||||
.to_string(),
|
||||
};
|
||||
if let Err(e) = self.sqlite.create_session(&session_row) {
|
||||
tracing::warn!(
|
||||
target: "kebab-app",
|
||||
error = %e,
|
||||
session_id = %session_id,
|
||||
"ask_with_session: create_session failed; continuing — turn append will surface a more useful error"
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
// Append the new turn. Failure is logged but does NOT mask
|
||||
// the answer — the user still gets their response, the
|
||||
// operator sees the persistence error in the warn log.
|
||||
let turn_id = format!(
|
||||
"{:032x}",
|
||||
blake3_truncate(&format!("{session_id}:{next_index}")),
|
||||
);
|
||||
let turn_row = ChatTurnRow {
|
||||
turn_id,
|
||||
session_id: session_id.to_string(),
|
||||
turn_index: next_index,
|
||||
question: query.to_string(),
|
||||
answer: answer.answer.clone(),
|
||||
citations_json: serde_json::to_string(&answer.citations)
|
||||
.unwrap_or_else(|_| "[]".to_string()),
|
||||
created_at: now_unix,
|
||||
};
|
||||
if let Err(e) = self.sqlite.append_turn(&turn_row) {
|
||||
tracing::warn!(
|
||||
target: "kebab-app",
|
||||
error = %e,
|
||||
session_id = %session_id,
|
||||
turn_index = next_index,
|
||||
"ask_with_session: append_turn failed; answer returned regardless"
|
||||
);
|
||||
}
|
||||
|
||||
Ok(answer)
|
||||
}
|
||||
|
||||
/// Returns `true` when the workspace has embeddings turned off
|
||||
/// (`provider = "none"` or `dimensions = 0`). Lexical-only mode.
|
||||
pub(crate) fn embeddings_disabled(&self) -> bool {
|
||||
@@ -834,28 +617,48 @@ impl App {
|
||||
if let Some(e) = self.embedder.get() {
|
||||
return Ok(Some(e.clone()));
|
||||
}
|
||||
// Provider branch (Track 1 spec §3 + arctic-embedder spec). The
|
||||
// `embeddings_disabled()` check above already handled `"none"`; here we
|
||||
// route the live providers. `fastembed`/`onnx`/(empty) keep the default
|
||||
// onnxruntime path (vectors unchanged — `embedding_version` is
|
||||
// preserved); `candle` selects the pure-Rust NUMA-safe backend (e5 or
|
||||
// arctic via its model registry); `ollama` offloads to a remote
|
||||
// `/api/embed` daemon.
|
||||
// Provider branch (arctic-embedder spec). The `embeddings_disabled()`
|
||||
// check above already handled `"none"`; here we route the live
|
||||
// providers. `fastembed`/`onnx`/(empty) keep the default onnxruntime
|
||||
// path (vectors unchanged — `embedding_version` is preserved); `ollama`
|
||||
// offloads to a remote `/api/embed` daemon.
|
||||
let provider = self.config.models.embedding.provider.as_str();
|
||||
let emb: Arc<dyn Embedder + Send + Sync> = match provider {
|
||||
"fastembed" | "onnx" | "" => Arc::new(
|
||||
FastembedEmbedder::new(&self.config).context("kb-app: load FastembedEmbedder")?,
|
||||
),
|
||||
"candle" => Arc::new(
|
||||
CandleEmbedder::new(&self.config).context("kb-app: load CandleEmbedder")?,
|
||||
),
|
||||
"ollama" => Arc::new(
|
||||
OllamaEmbedder::new(&self.config).context("kb-app: load OllamaEmbedder")?,
|
||||
),
|
||||
"fastembed" | "onnx" | "" => {
|
||||
// Resolve `{data_dir}/models/fastembed/` here so the
|
||||
// embedder constructor only takes the `[models.embedding]`
|
||||
// slice + the final cache dir.
|
||||
let data_dir = kebab_config::expand_path(&self.config.storage.data_dir, "");
|
||||
let model_dir = kebab_config::expand_path(
|
||||
&self.config.storage.model_dir,
|
||||
&data_dir.to_string_lossy(),
|
||||
);
|
||||
let cache_dir = model_dir.join(FASTEMBED_CACHE_SUBDIR);
|
||||
Arc::new(
|
||||
FastembedEmbedder::new(&self.config.models.embedding, &cache_dir)
|
||||
.context("kb-app: load FastembedEmbedder")?,
|
||||
)
|
||||
}
|
||||
"ollama" => {
|
||||
// Resolve the endpoint here: `models.embedding.endpoint`
|
||||
// → fallback `models.llm.endpoint`.
|
||||
let endpoint = self
|
||||
.config
|
||||
.models
|
||||
.embedding
|
||||
.endpoint
|
||||
.clone()
|
||||
.filter(|e| !e.is_empty())
|
||||
.unwrap_or_else(|| self.config.models.llm.endpoint.clone());
|
||||
Arc::new(
|
||||
OllamaEmbedder::new(&self.config.models.embedding, endpoint)
|
||||
.context("kb-app: load OllamaEmbedder")?,
|
||||
)
|
||||
}
|
||||
other => {
|
||||
return Err(anyhow!(
|
||||
"kb-app: unknown embedding provider {other:?}; expected one of \
|
||||
`fastembed` (default), `candle`, `ollama`, or `none` (lexical-only)"
|
||||
`fastembed` (default), `ollama`, or `none` (lexical-only)"
|
||||
));
|
||||
}
|
||||
};
|
||||
@@ -876,7 +679,7 @@ impl App {
|
||||
return Ok(Some(v.clone()));
|
||||
}
|
||||
let store = Arc::new(
|
||||
LanceVectorStore::new(&self.config, self.sqlite.clone())
|
||||
LanceVectorStore::new(&self.config.storage, self.sqlite.clone())
|
||||
.context("kb-app: open LanceVectorStore")?,
|
||||
);
|
||||
let _ = self.vector.set(store.clone());
|
||||
@@ -898,48 +701,6 @@ impl App {
|
||||
Ok(self.llm.get().cloned().unwrap_or(llm))
|
||||
}
|
||||
|
||||
/// p9-fb-19: build a `SearchCacheKey` for `query`. For lexical
|
||||
/// mode the embedding_version slot is left empty (no embedder
|
||||
/// identity contributes to the result). For vector / hybrid
|
||||
/// modes the embedder is built (cold-start) so the version
|
||||
/// label can be read; that's the cost the cache exists to
|
||||
/// amortize over the next few identical queries.
|
||||
fn build_cache_key(&self, query: &SearchQuery) -> Result<SearchCacheKey> {
|
||||
let embedding_version = match query.mode {
|
||||
SearchMode::Lexical => String::new(),
|
||||
SearchMode::Vector | SearchMode::Hybrid => {
|
||||
let emb = self.embedder()?.ok_or_else(|| {
|
||||
anyhow!(
|
||||
"embeddings disabled; vector / hybrid search require an \
|
||||
embedder — switch to --mode lexical or enable a provider"
|
||||
)
|
||||
})?;
|
||||
vector_index_version(emb.as_ref()).0
|
||||
}
|
||||
};
|
||||
Ok(SearchCacheKey {
|
||||
query_norm: SearchCacheKey::normalize_query(&query.text),
|
||||
mode: query.mode,
|
||||
k: u32::try_from(query.k).unwrap_or(u32::MAX),
|
||||
snippet_chars: u32::try_from(self.config.search.snippet_chars).unwrap_or(u32::MAX),
|
||||
embedding_version,
|
||||
chunker_version: self.config.ingest.chunking.chunker_version.clone(),
|
||||
corpus_revision: self.sqlite.corpus_revision(),
|
||||
})
|
||||
}
|
||||
|
||||
/// p9-fb-19: clear the in-process search cache. Useful for tests
|
||||
/// and for explicit user actions (e.g. a future `kebab cache
|
||||
/// clear` admin command). No-op when the cache is disabled.
|
||||
pub fn clear_search_cache(&self) {
|
||||
if let Some(cache) = self.search_cache.as_ref() {
|
||||
let mut guard = cache
|
||||
.lock()
|
||||
.unwrap_or_else(std::sync::PoisonError::into_inner);
|
||||
guard.clear();
|
||||
}
|
||||
}
|
||||
|
||||
/// p10-1A-2 Task 8b: back-fill `SearchHit.repo` from the originating
|
||||
/// document's `Metadata.repo` for every hit whose `repo` field is
|
||||
/// currently `None`. The search layer (kebab-search) constructs hits
|
||||
@@ -1059,33 +820,6 @@ fn vector_index_version(embedder: &dyn Embedder) -> IndexVersion {
|
||||
))
|
||||
}
|
||||
|
||||
/// p9-fb-18: derive a chat-session title from the first question.
|
||||
/// Trim, NFC, take first ~40 chars. Always non-empty (falls back
|
||||
/// to `"untitled"`) — same defensive shape as kebab-normalize's
|
||||
/// derive_title.
|
||||
fn first_question_title(question: &str) -> String {
|
||||
use unicode_normalization::UnicodeNormalization;
|
||||
let nfc: String = question.trim().nfc().collect();
|
||||
let truncated: String = nfc.chars().take(40).collect();
|
||||
if truncated.is_empty() {
|
||||
"untitled".to_string()
|
||||
} else {
|
||||
truncated
|
||||
}
|
||||
}
|
||||
|
||||
/// p9-fb-18: 32-hex `turn_id` derived from session_id + turn_index.
|
||||
/// blake3 hash truncated to first 16 bytes; format as 32-char lowercase
|
||||
/// hex so it slots into the `chat_turns.turn_id` column without
|
||||
/// collision concerns under any realistic per-session turn count.
|
||||
fn blake3_truncate(input: &str) -> u128 {
|
||||
let hash = blake3::hash(input.as_bytes());
|
||||
let bytes = hash.as_bytes();
|
||||
let mut buf = [0u8; 16];
|
||||
buf.copy_from_slice(&bytes[..16]);
|
||||
u128::from_be_bytes(buf)
|
||||
}
|
||||
|
||||
/// p9-fb-34: trim `s` to at most `n` Unicode scalar chars. Cheap
|
||||
/// alternative to a `.chars().take(n).collect::<String>()` pattern;
|
||||
/// reserves capacity proportional to UTF-8 worst case (4 bytes / char)
|
||||
@@ -1346,49 +1080,6 @@ impl App {
|
||||
}
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
/// p9-fb-18: title trims, NFC-normalizes, caps at 40 chars.
|
||||
#[test]
|
||||
fn first_question_title_trims_and_caps() {
|
||||
assert_eq!(first_question_title(" hello "), "hello");
|
||||
let long = "a".repeat(100);
|
||||
assert_eq!(first_question_title(&long).chars().count(), 40);
|
||||
}
|
||||
|
||||
/// p9-fb-18: empty / whitespace-only question falls back to
|
||||
/// `"untitled"` (never returns empty).
|
||||
#[test]
|
||||
fn first_question_title_falls_back_to_untitled() {
|
||||
assert_eq!(first_question_title(""), "untitled");
|
||||
assert_eq!(first_question_title(" "), "untitled");
|
||||
assert_eq!(first_question_title("\t\n"), "untitled");
|
||||
}
|
||||
|
||||
/// p9-fb-18: korean NFD → NFC.
|
||||
#[test]
|
||||
fn first_question_title_nfc_normalizes_korean() {
|
||||
let nfd = "\u{1100}\u{1161}".to_string(); // 가 (NFD)
|
||||
let title = first_question_title(&nfd);
|
||||
assert_eq!(title, "\u{AC00}", "expected NFC composed form");
|
||||
}
|
||||
|
||||
/// p9-fb-18: blake3_truncate is deterministic and differs across
|
||||
/// distinct inputs.
|
||||
#[test]
|
||||
fn blake3_truncate_deterministic_and_distinct() {
|
||||
let a = blake3_truncate("session-x:0");
|
||||
let b = blake3_truncate("session-x:0");
|
||||
let c = blake3_truncate("session-x:1");
|
||||
let d = blake3_truncate("session-y:0");
|
||||
assert_eq!(a, b, "same input → same hash");
|
||||
assert_ne!(a, c, "different turn_index → different hash");
|
||||
assert_ne!(a, d, "different session_id → different hash");
|
||||
}
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests_trace {
|
||||
use super::*;
|
||||
@@ -1399,7 +1090,7 @@ mod tests_trace {
|
||||
let mut cfg = kebab_config::Config::defaults();
|
||||
cfg.storage.data_dir = dir.path().to_string_lossy().into_owned();
|
||||
// Bring up migrations.
|
||||
let store = kebab_store_sqlite::SqliteStore::open(&cfg).unwrap();
|
||||
let store = kebab_store_sqlite::SqliteStore::open(&cfg.storage).unwrap();
|
||||
store.run_migrations().unwrap();
|
||||
drop(store);
|
||||
let app = App::open_with_config(cfg).unwrap();
|
||||
@@ -1451,7 +1142,7 @@ mod tests_trace {
|
||||
/// are `pub(crate)` — integration tests cannot reach them.
|
||||
///
|
||||
/// Spec §5.1 + plan §2 Step 10 — 3 test class:
|
||||
/// 1. registry length = 11 (image + pdf + 9 AST).
|
||||
/// 1. registry length = 12 (markdown + image + pdf + 9 AST).
|
||||
/// 2. mutually-exclusive `supports()` grid over 16 sample MediaTypes.
|
||||
/// 3. `extract_for` returns `Err("no Extractor ...")` for registry-NOT-cover
|
||||
/// MediaType (Audio).
|
||||
@@ -1467,28 +1158,26 @@ mod tests_extractor_dispatch {
|
||||
let mut cfg = kebab_config::Config::defaults();
|
||||
cfg.storage.data_dir = dir.path().to_string_lossy().into_owned();
|
||||
// Bring up migrations.
|
||||
let store = kebab_store_sqlite::SqliteStore::open(&cfg).unwrap();
|
||||
let store = kebab_store_sqlite::SqliteStore::open(&cfg.storage).unwrap();
|
||||
store.run_migrations().unwrap();
|
||||
drop(store);
|
||||
let app = App::open_with_config(cfg).unwrap();
|
||||
(dir, app)
|
||||
}
|
||||
|
||||
/// Registry length invariant: 11 Extractor (image + pdf + 9 AST).
|
||||
/// Markdown is NOT registered (free-function path — defer to a
|
||||
/// separate PR per spec §3.4).
|
||||
/// Registry length invariant: 12 Extractor (markdown + image + pdf +
|
||||
/// 9 AST). Markdown 합류로 모든 media 가 `extract_for` 경유로 통일됨.
|
||||
#[test]
|
||||
fn registry_has_eleven_extractors() {
|
||||
fn registry_has_twelve_extractors() {
|
||||
let (_dir, app) = open_app_with_temp_dir();
|
||||
assert_eq!(
|
||||
app.extractors.len(),
|
||||
11,
|
||||
"registry must hold 11 Extractors (image + pdf + 9 AST). \
|
||||
markdown 은 별 PR."
|
||||
12,
|
||||
"registry must hold 12 Extractors (markdown + image + pdf + 9 AST)."
|
||||
);
|
||||
}
|
||||
|
||||
/// 11 Extractor 의 `supports()` 가 16 sample MediaType 에 대해
|
||||
/// 12 Extractor 의 `supports()` 가 16 sample MediaType 에 대해
|
||||
/// mutually exclusive — 어떤 두 Extractor 도 동일 MediaType 에
|
||||
/// 대해 true 반환 안 됨.
|
||||
#[test]
|
||||
@@ -1559,6 +1248,8 @@ mod tests_extractor_dispatch {
|
||||
asset: &asset,
|
||||
workspace_root: &workspace_root,
|
||||
config: &cfg,
|
||||
source_id: None,
|
||||
source_trust: None,
|
||||
};
|
||||
let result = app.extract_for(&MediaType::Audio(AudioType::Wav), &ctx, &[]);
|
||||
assert!(result.is_err(), "Audio 는 registry 미포함 → Err 기대");
|
||||
|
||||
@@ -262,7 +262,7 @@ mod tests {
|
||||
let mut cfg = kebab_config::Config::defaults();
|
||||
cfg.storage.data_dir = dir.path().to_string_lossy().into_owned();
|
||||
// Bring up migrations so SqliteStore::open_existing succeeds inside App::open.
|
||||
let store = kebab_store_sqlite::SqliteStore::open(&cfg).unwrap();
|
||||
let store = kebab_store_sqlite::SqliteStore::open(&cfg.storage).unwrap();
|
||||
store.run_migrations().unwrap();
|
||||
drop(store);
|
||||
// Leak the tempdir into a static — tests are short-lived; not worth threading.
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
//! Typed signal re-exports + new signals introduced by fb-27.
|
||||
//!
|
||||
//! kebab-cli (and future kebab-tui / kebab-desktop) downcast on these to
|
||||
//! kebab-cli (and future kebab-desktop) downcast on these to
|
||||
//! build `error.v1` wire records. The existing signals
|
||||
//! (`RefusalSignal`, `NoHitSignal`, `DoctorUnhealthy`) live in
|
||||
//! `doctor_signal.rs` — leave those unchanged and re-export via this
|
||||
|
||||
3867
crates/kebab-app/src/ingest.rs
Normal file
3867
crates/kebab-app/src/ingest.rs
Normal file
File diff suppressed because it is too large
Load Diff
@@ -1,4 +1,4 @@
|
||||
//! Streaming progress events for `ingest_with_config_progress`.
|
||||
//! Streaming progress events for `ingest_with_config` (via `IngestOpts::progress`).
|
||||
//!
|
||||
//! The facade emits one [`IngestEvent`] per step boundary into an
|
||||
//! optional `mpsc::Sender<IngestEvent>` injected by the caller. CLI
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -139,7 +139,7 @@ pub fn enumerate_orphans(cfg: &Config) -> Result<Vec<WorkspacePath>> {
|
||||
use kebab_core::SourceScope;
|
||||
use kebab_source_fs::FsSourceConnector;
|
||||
|
||||
let store = kebab_store_sqlite::SqliteStore::open(cfg)
|
||||
let store = kebab_store_sqlite::SqliteStore::open(&cfg.storage)
|
||||
.context("enumerate_orphans: open SqliteStore")?;
|
||||
|
||||
let stored = store
|
||||
@@ -237,7 +237,7 @@ fn execute_orphans_only(cfg: &Config) -> Result<ResetReport> {
|
||||
}
|
||||
|
||||
let store = std::sync::Arc::new(
|
||||
kebab_store_sqlite::SqliteStore::open(cfg)
|
||||
kebab_store_sqlite::SqliteStore::open(&cfg.storage)
|
||||
.context("execute_orphans_only: open SqliteStore")?,
|
||||
);
|
||||
|
||||
@@ -296,7 +296,7 @@ fn open_vector_store_if_configured(
|
||||
if cfg.models.embedding.provider == "none" || cfg.models.embedding.dimensions == 0 {
|
||||
return Ok(None);
|
||||
}
|
||||
match kebab_store_vector::LanceVectorStore::new(cfg, store) {
|
||||
match kebab_store_vector::LanceVectorStore::new(&cfg.storage, store) {
|
||||
Ok(vs) => Ok(Some(vs)),
|
||||
Err(e) => {
|
||||
tracing::warn!(
|
||||
@@ -320,7 +320,7 @@ fn truncate_embeddings(cfg: &Config) -> Result<u64> {
|
||||
if !sqlite_path.exists() {
|
||||
return Ok(0);
|
||||
}
|
||||
let store = kebab_store_sqlite::SqliteStore::open(cfg)
|
||||
let store = kebab_store_sqlite::SqliteStore::open(&cfg.storage)
|
||||
.context("open SqliteStore for truncate_embedding_records")?;
|
||||
store.truncate_embedding_records()
|
||||
}
|
||||
|
||||
@@ -151,8 +151,8 @@ fn capabilities_snapshot() -> Capabilities {
|
||||
json_mode: true,
|
||||
ingest_progress: true,
|
||||
ingest_cancellation: true,
|
||||
rag_multi_turn: true,
|
||||
search_cache: true,
|
||||
rag_multi_turn: false,
|
||||
search_cache: false,
|
||||
incremental_ingest: true,
|
||||
streaming_ask: true,
|
||||
http_daemon: false,
|
||||
@@ -263,7 +263,7 @@ mod tests_stats_ext {
|
||||
let mut cfg = kebab_config::Config::defaults();
|
||||
cfg.storage.data_dir = dir.path().to_string_lossy().into_owned();
|
||||
// Bring up migrations so the sqlite file is created.
|
||||
let store = kebab_store_sqlite::SqliteStore::open(&cfg).unwrap();
|
||||
let store = kebab_store_sqlite::SqliteStore::open(&cfg.storage).unwrap();
|
||||
store.run_migrations().unwrap();
|
||||
drop(store);
|
||||
|
||||
|
||||
@@ -21,7 +21,7 @@ use common::TestEnv;
|
||||
#[ignore = "requires real Ollama on 127.0.0.1:11434"]
|
||||
fn ask_lexical_smoke() {
|
||||
let env = TestEnv::lexical_only();
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), true).unwrap();
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts { summary_only: true, ..Default::default() }).unwrap();
|
||||
|
||||
let opts = kebab_app::AskOpts {
|
||||
k: 5,
|
||||
@@ -30,9 +30,6 @@ fn ask_lexical_smoke() {
|
||||
temperature: Some(0.0),
|
||||
seed: Some(0),
|
||||
stream_sink: None,
|
||||
history: Vec::new(),
|
||||
conversation_id: None,
|
||||
turn_index: None,
|
||||
multi_hop: false,
|
||||
};
|
||||
// The fixture workspace contains "ownership" content; the model's
|
||||
|
||||
@@ -29,7 +29,7 @@ fn rust_file_ingests_and_searches_as_code_citation() {
|
||||
)
|
||||
.unwrap();
|
||||
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), false)
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default())
|
||||
.expect("ingest must succeed");
|
||||
|
||||
assert_eq!(report.errors, 0, "no errors expected: {report:?}");
|
||||
@@ -128,7 +128,7 @@ fn rust_code_search_hit_has_repo() {
|
||||
)
|
||||
.unwrap();
|
||||
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), false)
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default())
|
||||
.expect("ingest must succeed");
|
||||
assert_eq!(report.errors, 0, "no ingest errors: {report:?}");
|
||||
|
||||
@@ -176,7 +176,7 @@ fn python_file_ingests_and_searches_as_code_citation() {
|
||||
)
|
||||
.unwrap();
|
||||
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), false)
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default())
|
||||
.expect("ingest must succeed");
|
||||
|
||||
assert!(report.new >= 1, "python file ingested: {report:?}");
|
||||
@@ -254,7 +254,7 @@ fn typescript_file_ingests_and_searches_as_code_citation() {
|
||||
)
|
||||
.unwrap();
|
||||
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), false)
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default())
|
||||
.expect("ingest must succeed");
|
||||
|
||||
assert!(report.new >= 1, "ts file ingested: {report:?}");
|
||||
@@ -332,7 +332,7 @@ fn javascript_file_ingests_and_searches_as_code_citation() {
|
||||
)
|
||||
.unwrap();
|
||||
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), false)
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default())
|
||||
.expect("ingest must succeed");
|
||||
|
||||
assert!(report.new >= 1, "js file ingested: {report:?}");
|
||||
@@ -410,7 +410,7 @@ fn go_file_ingests_and_searches_as_code_citation() {
|
||||
)
|
||||
.unwrap();
|
||||
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), false)
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default())
|
||||
.expect("ingest must succeed");
|
||||
assert_eq!(report.errors, 0);
|
||||
assert!(report.new >= 1);
|
||||
@@ -483,7 +483,7 @@ fn java_file_ingests_and_searches_as_code_citation() {
|
||||
)
|
||||
.unwrap();
|
||||
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), false)
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default())
|
||||
.expect("ingest must succeed");
|
||||
assert_eq!(report.errors, 0);
|
||||
assert!(report.new >= 1);
|
||||
@@ -560,7 +560,7 @@ fn kotlin_file_ingests_and_searches_as_code_citation() {
|
||||
)
|
||||
.unwrap();
|
||||
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), false)
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default())
|
||||
.expect("ingest must succeed");
|
||||
assert_eq!(report.errors, 0);
|
||||
assert!(report.new >= 1);
|
||||
@@ -635,7 +635,7 @@ fn tier2_k8s_yaml_ingest_searchable() {
|
||||
)
|
||||
.unwrap();
|
||||
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), false)
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default())
|
||||
.expect("ingest must succeed");
|
||||
assert_eq!(report.errors, 0, "no ingest errors: {report:?}");
|
||||
assert!(report.new >= 1, "yaml file ingested: {report:?}");
|
||||
@@ -720,7 +720,7 @@ fn tier2_dockerfile_ingest_searchable() {
|
||||
)
|
||||
.unwrap();
|
||||
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), false)
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default())
|
||||
.expect("ingest must succeed");
|
||||
assert_eq!(report.errors, 0, "no ingest errors: {report:?}");
|
||||
assert!(report.new >= 1, "Dockerfile ingested: {report:?}");
|
||||
@@ -805,7 +805,7 @@ fn tier2_cargo_toml_ingest_searchable() {
|
||||
)
|
||||
.unwrap();
|
||||
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), false)
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default())
|
||||
.expect("ingest must succeed");
|
||||
assert_eq!(report.errors, 0, "no ingest errors: {report:?}");
|
||||
assert!(report.new >= 1, "Cargo.toml ingested: {report:?}");
|
||||
@@ -890,7 +890,7 @@ fn tier3_shell_ingest_searchable() {
|
||||
)
|
||||
.unwrap();
|
||||
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), false)
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default())
|
||||
.expect("ingest must succeed");
|
||||
assert_eq!(report.errors, 0, "no ingest errors: {report:?}");
|
||||
assert!(report.new >= 1, "shell file ingested: {report:?}");
|
||||
@@ -979,7 +979,7 @@ fn tier3_yaml_fallback_picks_up_non_k8s_yaml() {
|
||||
)
|
||||
.unwrap();
|
||||
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), false)
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default())
|
||||
.expect("ingest must succeed");
|
||||
assert_eq!(report.errors, 0, "no ingest errors: {report:?}");
|
||||
assert!(
|
||||
@@ -1063,7 +1063,7 @@ fn rust_file_re_ingest_is_unchanged() {
|
||||
|
||||
std::fs::write(env.workspace_root.join("stable.rs"), "pub fn noop() {}\n").unwrap();
|
||||
|
||||
let r1 = kebab_app::ingest_with_config(env.config.clone(), env.scope(), false).unwrap();
|
||||
let r1 = kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default()).unwrap();
|
||||
let item1 = r1
|
||||
.items
|
||||
.as_ref()
|
||||
@@ -1074,7 +1074,7 @@ fn rust_file_re_ingest_is_unchanged() {
|
||||
.unwrap();
|
||||
assert_eq!(item1.kind, IngestItemKind::New);
|
||||
|
||||
let r2 = kebab_app::ingest_with_config(env.config.clone(), env.scope(), false).unwrap();
|
||||
let r2 = kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default()).unwrap();
|
||||
let item2 = r2
|
||||
.items
|
||||
.unwrap()
|
||||
@@ -1105,7 +1105,7 @@ fn tier3_yaml_fallback_reingest_is_unchanged() {
|
||||
)
|
||||
.unwrap();
|
||||
|
||||
let report1 = kebab_app::ingest_with_config(env.config.clone(), env.scope(), false)
|
||||
let report1 = kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default())
|
||||
.expect("first ingest");
|
||||
let item1 = report1
|
||||
.items
|
||||
@@ -1125,7 +1125,7 @@ fn tier3_yaml_fallback_reingest_is_unchanged() {
|
||||
"first ingest must use Tier 3 fallback chunker"
|
||||
);
|
||||
|
||||
let report2 = kebab_app::ingest_with_config(env.config.clone(), env.scope(), false)
|
||||
let report2 = kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default())
|
||||
.expect("second ingest");
|
||||
let item2 = report2
|
||||
.items
|
||||
@@ -1155,7 +1155,7 @@ fn tier1_c_ingest_searchable() {
|
||||
)
|
||||
.unwrap();
|
||||
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), false)
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default())
|
||||
.expect("ingest must succeed");
|
||||
assert_eq!(report.errors, 0, "no ingest errors: {report:?}");
|
||||
assert!(report.new >= 1, "c file ingested: {report:?}");
|
||||
@@ -1241,7 +1241,7 @@ fn tier1_cpp_ingest_searchable() {
|
||||
)
|
||||
.unwrap();
|
||||
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), false)
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default())
|
||||
.expect("ingest must succeed");
|
||||
assert_eq!(report.errors, 0, "no ingest errors: {report:?}");
|
||||
assert!(report.new >= 1, "cpp file ingested: {report:?}");
|
||||
@@ -1333,7 +1333,7 @@ fn tier2_k8s_multi_resource_yaml_ingests_without_collision() {
|
||||
)
|
||||
.unwrap();
|
||||
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), false)
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default())
|
||||
.expect("ingest must succeed");
|
||||
|
||||
// The bug: this would land in report with an error + UNIQUE constraint message.
|
||||
@@ -1389,7 +1389,7 @@ fn tier3_shell_reingest_is_unchanged() {
|
||||
)
|
||||
.unwrap();
|
||||
|
||||
let report1 = kebab_app::ingest_with_config(env.config.clone(), env.scope(), false)
|
||||
let report1 = kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default())
|
||||
.expect("first ingest");
|
||||
let item1 = report1
|
||||
.items
|
||||
@@ -1404,7 +1404,7 @@ fn tier3_shell_reingest_is_unchanged() {
|
||||
item1.kind
|
||||
);
|
||||
|
||||
let report2 = kebab_app::ingest_with_config(env.config.clone(), env.scope(), false)
|
||||
let report2 = kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default())
|
||||
.expect("second ingest");
|
||||
let item2 = report2
|
||||
.items
|
||||
|
||||
@@ -107,7 +107,7 @@ pub fn ingest_md(env: &TestEnv, relative_path: &str, content: &str) {
|
||||
std::fs::create_dir_all(parent).expect("create parent dirs");
|
||||
}
|
||||
std::fs::write(&path, content).expect("write workspace file");
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), true)
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts { summary_only: true, ..Default::default() })
|
||||
.expect("ingest_with_config");
|
||||
}
|
||||
|
||||
|
||||
@@ -15,7 +15,7 @@ mod common;
|
||||
|
||||
use common::TestEnv;
|
||||
|
||||
use kebab_app::{IngestOpts, ingest_with_config, ingest_with_config_opts};
|
||||
use kebab_app::{IngestOpts, ingest_with_config};
|
||||
use kebab_core::IngestItemKind;
|
||||
|
||||
/// Seed a workspace with a markdown + a rust file so both the markdown and
|
||||
@@ -26,7 +26,7 @@ fn seed_and_first_ingest(env: &TestEnv) -> kebab_core::IngestReport {
|
||||
"/// adds two integers\npub fn add(a: i32, b: i32) -> i32 {\n a + b\n}\n",
|
||||
)
|
||||
.unwrap();
|
||||
let first = ingest_with_config(env.config.clone(), env.scope(), false).expect("first ingest");
|
||||
let first = ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default()).expect("first ingest");
|
||||
assert_eq!(first.errors, 0, "first ingest must not error: {first:?}");
|
||||
assert!(first.new >= 1, "first ingest creates docs: {first:?}");
|
||||
assert_eq!(first.unchanged, 0, "first ingest has no unchanged: {first:?}");
|
||||
@@ -34,7 +34,7 @@ fn seed_and_first_ingest(env: &TestEnv) -> kebab_core::IngestReport {
|
||||
}
|
||||
|
||||
fn reingest(env: &TestEnv) -> kebab_core::IngestReport {
|
||||
ingest_with_config_opts(env.config.clone(), env.scope(), false, IngestOpts::default())
|
||||
ingest_with_config(env.config.clone(), env.scope(), IngestOpts::default())
|
||||
.expect("re-ingest")
|
||||
}
|
||||
|
||||
|
||||
@@ -16,14 +16,13 @@
|
||||
mod common;
|
||||
|
||||
use common::TestEnv;
|
||||
use kebab_app::IngestOpts;
|
||||
use kebab_app::ingest_with_config_opts;
|
||||
use kebab_app::{IngestOpts, ingest_with_config};
|
||||
use kebab_core::{DocFilter, DocumentStore, SearchMode, SearchQuery, SourceScope};
|
||||
|
||||
/// Helper: open the store via `TestEnv` and run `list_documents`.
|
||||
fn list_doc_paths(env: &TestEnv) -> Vec<String> {
|
||||
use kebab_store_sqlite::SqliteStore;
|
||||
let store = SqliteStore::open(&env.config).unwrap();
|
||||
let store = SqliteStore::open(&env.config.storage).unwrap();
|
||||
store.run_migrations().unwrap();
|
||||
store
|
||||
.list_documents(&DocFilter::default())
|
||||
@@ -44,10 +43,9 @@ fn file_deletion_auto_purge() {
|
||||
std::fs::write(&b_path, "// file b\nfn bravo() {}\n").unwrap();
|
||||
|
||||
// First ingest — both must be New.
|
||||
let first = ingest_with_config_opts(
|
||||
let first = ingest_with_config(
|
||||
env.config.clone(),
|
||||
env.scope(),
|
||||
false,
|
||||
IngestOpts::default(),
|
||||
)
|
||||
.expect("first ingest must succeed");
|
||||
@@ -64,10 +62,9 @@ fn file_deletion_auto_purge() {
|
||||
std::fs::remove_file(&b_path).expect("remove b.rs");
|
||||
|
||||
// Second ingest — scanned count drops by 1; b.rs should be purged.
|
||||
let second = ingest_with_config_opts(
|
||||
let second = ingest_with_config(
|
||||
env.config.clone(),
|
||||
env.scope(),
|
||||
false,
|
||||
IngestOpts::default(),
|
||||
)
|
||||
.expect("second ingest must succeed");
|
||||
@@ -126,7 +123,7 @@ fn include_scope_narrowing_does_not_purge() {
|
||||
exclude: env.config.workspace.exclude.clone(),
|
||||
};
|
||||
let first =
|
||||
ingest_with_config_opts(env.config.clone(), wide_scope, false, IngestOpts::default())
|
||||
ingest_with_config(env.config.clone(), wide_scope, IngestOpts::default())
|
||||
.expect("first ingest (wide) must succeed");
|
||||
assert!(first.new >= 2, "expected at least 2 new docs: {first:?}");
|
||||
assert_eq!(
|
||||
@@ -141,10 +138,9 @@ fn include_scope_narrowing_does_not_purge() {
|
||||
include: vec!["a_narrow.rs".to_string()],
|
||||
exclude: env.config.workspace.exclude.clone(),
|
||||
};
|
||||
let second = ingest_with_config_opts(
|
||||
let second = ingest_with_config(
|
||||
env.config.clone(),
|
||||
narrow_scope,
|
||||
false,
|
||||
IngestOpts::default(),
|
||||
)
|
||||
.expect("second ingest (narrow) must succeed");
|
||||
|
||||
@@ -71,7 +71,7 @@ async fn ingest_image_with_ocr_produces_chunk_containing_ocr_text() {
|
||||
let env_scope = env.scope();
|
||||
|
||||
let report = spawn_blocking(move || {
|
||||
kebab_app::ingest_with_config(cfg_clone, env_scope, false)
|
||||
kebab_app::ingest_with_config(cfg_clone, env_scope, kebab_app::IngestOpts::default())
|
||||
.expect("image ingest must succeed")
|
||||
})
|
||||
.await
|
||||
@@ -167,7 +167,7 @@ async fn ingest_image_with_ocr_and_caption_populates_both_fields() {
|
||||
let cfg_clone = cfg.clone();
|
||||
let scope = env.scope();
|
||||
let report = spawn_blocking(move || {
|
||||
kebab_app::ingest_with_config(cfg_clone, scope, false)
|
||||
kebab_app::ingest_with_config(cfg_clone, scope, kebab_app::IngestOpts::default())
|
||||
.expect("ingest must succeed with both OCR+caption")
|
||||
})
|
||||
.await
|
||||
@@ -212,7 +212,7 @@ async fn ocr_failure_indexes_asset_with_warning_no_error_counter() {
|
||||
let cfg_clone = cfg.clone();
|
||||
let scope = env.scope();
|
||||
let report = spawn_blocking(move || {
|
||||
kebab_app::ingest_with_config(cfg_clone, scope, false)
|
||||
kebab_app::ingest_with_config(cfg_clone, scope, kebab_app::IngestOpts::default())
|
||||
.expect("ingest does not abort on lenient OCR failure")
|
||||
})
|
||||
.await
|
||||
@@ -276,7 +276,7 @@ async fn image_indexed_with_filename_when_ocr_and_caption_disabled() {
|
||||
let cfg_clone = cfg.clone();
|
||||
let scope = env.scope();
|
||||
let report = spawn_blocking(move || {
|
||||
kebab_app::ingest_with_config(cfg_clone, scope, false).expect("ingest with no OCR/caption")
|
||||
kebab_app::ingest_with_config(cfg_clone, scope, kebab_app::IngestOpts::default()).expect("ingest with no OCR/caption")
|
||||
})
|
||||
.await
|
||||
.expect("task");
|
||||
@@ -340,7 +340,7 @@ async fn garbage_png_increments_errors_counter_exactly_once() {
|
||||
let cfg_clone = cfg.clone();
|
||||
let scope = env.scope();
|
||||
let report = spawn_blocking(move || {
|
||||
kebab_app::ingest_with_config(cfg_clone, scope, false)
|
||||
kebab_app::ingest_with_config(cfg_clone, scope, kebab_app::IngestOpts::default())
|
||||
.expect("ingest does not abort on per-asset failure")
|
||||
})
|
||||
.await
|
||||
@@ -399,10 +399,10 @@ async fn re_ingest_image_produces_unchanged_with_same_doc_id() {
|
||||
let scope1 = scope.clone();
|
||||
let scope2 = scope.clone();
|
||||
|
||||
let r1 = spawn_blocking(move || kebab_app::ingest_with_config(cfg1, scope1, false).unwrap())
|
||||
let r1 = spawn_blocking(move || kebab_app::ingest_with_config(cfg1, scope1, kebab_app::IngestOpts::default()).unwrap())
|
||||
.await
|
||||
.unwrap();
|
||||
let r2 = spawn_blocking(move || kebab_app::ingest_with_config(cfg2, scope2, false).unwrap())
|
||||
let r2 = spawn_blocking(move || kebab_app::ingest_with_config(cfg2, scope2, kebab_app::IngestOpts::default()).unwrap())
|
||||
.await
|
||||
.unwrap();
|
||||
|
||||
|
||||
@@ -12,7 +12,7 @@ mod common;
|
||||
|
||||
use common::TestEnv;
|
||||
|
||||
use kebab_app::{IngestOpts, ingest_with_config, ingest_with_config_opts};
|
||||
use kebab_app::{IngestOpts, ingest_with_config};
|
||||
|
||||
#[test]
|
||||
fn second_ingest_of_unchanged_corpus_marks_all_unchanged() {
|
||||
@@ -21,7 +21,7 @@ fn second_ingest_of_unchanged_corpus_marks_all_unchanged() {
|
||||
// First ingest — populates the DB. Use the legacy entry so the
|
||||
// assertions cover the "previously ingested" set without needing
|
||||
// IngestOpts::default() to behave identically.
|
||||
let first = ingest_with_config(env.config.clone(), env.scope(), false).unwrap();
|
||||
let first = ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default()).unwrap();
|
||||
assert_eq!(first.errors, 0, "first ingest must not error: {first:?}");
|
||||
assert!(
|
||||
first.new >= 1,
|
||||
@@ -36,13 +36,8 @@ fn second_ingest_of_unchanged_corpus_marks_all_unchanged() {
|
||||
|
||||
// Second ingest — same files, same versions → all assets must be
|
||||
// labelled Unchanged (no parse / chunk / embed re-work).
|
||||
let second = ingest_with_config_opts(
|
||||
env.config.clone(),
|
||||
env.scope(),
|
||||
false,
|
||||
IngestOpts::default(),
|
||||
)
|
||||
.unwrap();
|
||||
let second = ingest_with_config(env.config.clone(), env.scope(), IngestOpts::default())
|
||||
.unwrap();
|
||||
assert_eq!(
|
||||
second.scanned, scanned,
|
||||
"second scanned matches first: {second:?}"
|
||||
@@ -63,7 +58,7 @@ fn second_ingest_of_unchanged_corpus_marks_all_unchanged() {
|
||||
fn force_reingest_bypasses_skip() {
|
||||
let env = TestEnv::lexical_only();
|
||||
|
||||
let first = ingest_with_config(env.config.clone(), env.scope(), false).unwrap();
|
||||
let first = ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default()).unwrap();
|
||||
assert_eq!(first.errors, 0, "first ingest must not error: {first:?}");
|
||||
assert!(
|
||||
first.new >= 1,
|
||||
@@ -71,10 +66,9 @@ fn force_reingest_bypasses_skip() {
|
||||
);
|
||||
let scanned = first.scanned;
|
||||
|
||||
let second = ingest_with_config_opts(
|
||||
let second = ingest_with_config(
|
||||
env.config.clone(),
|
||||
env.scope(),
|
||||
false,
|
||||
IngestOpts {
|
||||
force_reingest: true,
|
||||
..Default::default()
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
//! Integration coverage for `ingest_with_config_cancellable`
|
||||
//! Integration coverage for cancellable ingest via `IngestOpts`
|
||||
//! (p9-fb-04). Asserts the §10 invariants:
|
||||
//!
|
||||
//! - Cancel set BEFORE the loop starts → no asset is processed.
|
||||
@@ -21,12 +21,15 @@ fn run_with(
|
||||
cancel: Arc<AtomicBool>,
|
||||
progress: Option<mpsc::Sender<IngestEvent>>,
|
||||
) -> kebab_core::IngestReport {
|
||||
kebab_app::ingest_with_config_cancellable(
|
||||
kebab_app::ingest_with_config(
|
||||
env.config.clone(),
|
||||
env.scope(),
|
||||
true,
|
||||
progress,
|
||||
Some(cancel),
|
||||
kebab_app::IngestOpts {
|
||||
progress,
|
||||
cancel: Some(cancel),
|
||||
summary_only: true,
|
||||
..Default::default()
|
||||
},
|
||||
)
|
||||
.unwrap()
|
||||
}
|
||||
@@ -89,7 +92,15 @@ fn cancel_mid_loop_after_first_asset_keeps_idempotent_resume() {
|
||||
assert!(report.new < 3, "loop should have broken: {report:?}");
|
||||
|
||||
// Idempotent re-ingest finishes the job.
|
||||
let r2 = kebab_app::ingest_with_config(env.config.clone(), env.scope(), true).unwrap();
|
||||
let r2 = kebab_app::ingest_with_config(
|
||||
env.config.clone(),
|
||||
env.scope(),
|
||||
kebab_app::IngestOpts {
|
||||
summary_only: true,
|
||||
..Default::default()
|
||||
},
|
||||
)
|
||||
.unwrap();
|
||||
assert_eq!(r2.scanned, 3, "re-scan: {r2:?}");
|
||||
// Total committed across both runs covers all 3 docs (some New
|
||||
// first run, rest New on second; or first run was 0 → all New on
|
||||
@@ -108,7 +119,15 @@ fn cancel_none_is_uncancellable_default() {
|
||||
let env = TestEnv::lexical_only();
|
||||
let (tx, rx) = mpsc::channel::<IngestEvent>();
|
||||
let report =
|
||||
kebab_app::ingest_with_config_progress(env.config.clone(), env.scope(), true, Some(tx))
|
||||
kebab_app::ingest_with_config(
|
||||
env.config.clone(),
|
||||
env.scope(),
|
||||
kebab_app::IngestOpts {
|
||||
progress: Some(tx),
|
||||
summary_only: true,
|
||||
..Default::default()
|
||||
},
|
||||
)
|
||||
.unwrap();
|
||||
assert_eq!(report.scanned, 3);
|
||||
assert_eq!(report.new, 3);
|
||||
|
||||
@@ -8,7 +8,7 @@ use common::TestEnv;
|
||||
#[test]
|
||||
fn ingest_then_list_inspects_round_trip() {
|
||||
let env = TestEnv::lexical_only();
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), false).unwrap();
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default()).unwrap();
|
||||
|
||||
// The fixture has 3 markdown files; first ingest should label them
|
||||
// all as New.
|
||||
@@ -42,10 +42,10 @@ fn ingest_then_list_inspects_round_trip() {
|
||||
fn ingest_idempotent_on_second_run() {
|
||||
let env = TestEnv::lexical_only();
|
||||
|
||||
let r1 = kebab_app::ingest_with_config(env.config.clone(), env.scope(), false).unwrap();
|
||||
let r1 = kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default()).unwrap();
|
||||
assert_eq!(r1.new, 3);
|
||||
|
||||
let r2 = kebab_app::ingest_with_config(env.config.clone(), env.scope(), false).unwrap();
|
||||
let r2 = kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default()).unwrap();
|
||||
// Same files re-ingested — p9-fb-23 task 7 introduced the early-skip
|
||||
// path: when checksum + parser/chunker/embedding versions all match,
|
||||
// the second run reports `Unchanged` rather than `Updated`. Pre-p9-fb-23
|
||||
@@ -66,7 +66,7 @@ fn ingest_idempotent_on_second_run() {
|
||||
#[test]
|
||||
fn ingest_summary_only_drops_items() {
|
||||
let env = TestEnv::lexical_only();
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), true).unwrap();
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts { summary_only: true, ..Default::default() }).unwrap();
|
||||
assert_eq!(report.scanned, 3);
|
||||
assert!(report.items.is_none(), "summary-only should null items");
|
||||
}
|
||||
@@ -78,7 +78,7 @@ fn ingest_records_ingest_runs_row_with_aggregate_counts() {
|
||||
// of every run. `summary_only=true` writes `items_json=NULL`; the
|
||||
// counts MUST still be present.
|
||||
let env = TestEnv::lexical_only();
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), true).unwrap();
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts { summary_only: true, ..Default::default() }).unwrap();
|
||||
assert_eq!(report.scanned, 3);
|
||||
|
||||
let db_path = std::path::PathBuf::from(&env.config.storage.data_dir).join("kebab.sqlite");
|
||||
@@ -130,7 +130,7 @@ fn ingest_provider_none_skips_lance() {
|
||||
// tree shape (no `<data_dir>/lancedb` directory, or no `*.lance`
|
||||
// tables under it).
|
||||
let env = TestEnv::lexical_only();
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), false).unwrap();
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default()).unwrap();
|
||||
assert_eq!(report.errors, 0, "lexical-only run must not error");
|
||||
assert_eq!(report.new, 3);
|
||||
|
||||
@@ -157,7 +157,7 @@ fn ingest_provider_none_skips_lance() {
|
||||
#[test]
|
||||
fn list_docs_filters_by_tags_any() {
|
||||
let env = TestEnv::lexical_only();
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), true).unwrap();
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts { summary_only: true, ..Default::default() }).unwrap();
|
||||
|
||||
let filter = kebab_core::DocFilter {
|
||||
tags_any: vec!["python".to_string()],
|
||||
@@ -205,16 +205,14 @@ fn inspect_chunk_not_found_returns_actionable_error() {
|
||||
assert!(msg.contains("not found"), "got: {msg}");
|
||||
}
|
||||
|
||||
/// p9-fb-23 task 6: `ingest_with_config_opts` with `IngestOpts::default()`
|
||||
/// must behave identically to `ingest_with_config` — first ingest reports
|
||||
/// all assets as new, no errors, no unchanged.
|
||||
/// p9-fb-23 task 6: `ingest_with_config` with `IngestOpts::default()`
|
||||
/// must report all assets as new, no errors, no unchanged on first ingest.
|
||||
#[test]
|
||||
fn ingest_with_config_opts_default_matches_legacy_behaviour() {
|
||||
let env = TestEnv::lexical_only();
|
||||
let report = kebab_app::ingest_with_config_opts(
|
||||
let report = kebab_app::ingest_with_config(
|
||||
env.config.clone(),
|
||||
env.scope(),
|
||||
false,
|
||||
kebab_app::IngestOpts::default(),
|
||||
)
|
||||
.unwrap();
|
||||
@@ -232,7 +230,7 @@ fn ingest_with_config_opts_default_matches_legacy_behaviour() {
|
||||
#[test]
|
||||
fn ingest_stamps_chunker_version_on_document() {
|
||||
let env = TestEnv::lexical_only();
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), false).unwrap();
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default()).unwrap();
|
||||
assert!(report.new >= 1, "expected at least one new doc: {report:?}");
|
||||
assert_eq!(report.errors, 0, "no errors expected: {report:?}");
|
||||
|
||||
|
||||
@@ -4,7 +4,7 @@
|
||||
|
||||
use std::path::PathBuf;
|
||||
|
||||
use kebab_app::{IngestOpts, ingest_with_config_opts};
|
||||
use kebab_app::{IngestOpts, ingest_with_config};
|
||||
use kebab_config::{Config, LoggingCfg};
|
||||
use kebab_core::SourceScope;
|
||||
use serde_json::Value;
|
||||
@@ -61,7 +61,7 @@ fn ingest_log_smoke() {
|
||||
};
|
||||
|
||||
// 3. Run ingest.
|
||||
ingest_with_config_opts(cfg, scope, false, IngestOpts::default())
|
||||
ingest_with_config(cfg, scope, IngestOpts::default())
|
||||
.expect("ingest should succeed");
|
||||
|
||||
// 4. Assert log file exists in log_dir.
|
||||
@@ -148,7 +148,7 @@ fn ingest_log_disabled_emits_no_file() {
|
||||
..Default::default()
|
||||
};
|
||||
|
||||
ingest_with_config_opts(cfg, scope, false, IngestOpts::default())
|
||||
ingest_with_config(cfg, scope, IngestOpts::default())
|
||||
.expect("ingest should succeed");
|
||||
|
||||
// log_dir should either not exist or contain 0 ingest-*.ndjson files.
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
//!
|
||||
//! Tests 1 and 2 require a live Ollama endpoint — `#[ignore]` by default.
|
||||
//! Manual invoke:
|
||||
//! KEBAB_PDF_OCR_ENDPOINT=http://192.168.0.47:11434 \
|
||||
//! KEBAB_OCR_ENDPOINT=http://192.168.0.47:11434 \
|
||||
//! cargo test -p kebab-app --test ingest_pdf_ocr_smoke --ignored -j 4
|
||||
//!
|
||||
//! Test 3 (cancel) uses a dummy endpoint + pre-set cancel — runs by default
|
||||
@@ -17,7 +17,8 @@ use std::sync::atomic::AtomicBool;
|
||||
use common::TestEnv;
|
||||
|
||||
fn ollama_endpoint() -> String {
|
||||
std::env::var("KEBAB_PDF_OCR_ENDPOINT").unwrap_or_else(|_| "http://localhost:11434".to_string())
|
||||
// v5: shared KEBAB_OCR_* env (manual harness reads it directly).
|
||||
std::env::var("KEBAB_OCR_ENDPOINT").unwrap_or_else(|_| "http://localhost:11434".to_string())
|
||||
}
|
||||
|
||||
fn make_ocr_env_real() -> TestEnv {
|
||||
@@ -43,7 +44,7 @@ fn ingest_with_mock_ocr_yields_pdf_ocr_summary() {
|
||||
let env = make_ocr_env_real();
|
||||
|
||||
let report =
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), false).expect("ingest");
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default()).expect("ingest");
|
||||
|
||||
assert!(report.new >= 1, "at least one PDF ingested: {report:?}");
|
||||
|
||||
@@ -71,7 +72,7 @@ fn ingest_with_mock_ocr_yields_pdf_ocr_summary() {
|
||||
fn ocr_text_indexed_and_searchable() {
|
||||
let env = make_ocr_env_real();
|
||||
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), false).expect("ingest");
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default()).expect("ingest");
|
||||
|
||||
// Search for a Korean morpheme expected to appear in qwen2.5vl:3b OCR
|
||||
// output of the PoC ground-truth page. "다음" is a high-frequency token
|
||||
@@ -104,12 +105,13 @@ fn ingest_with_cancel_aborts_mid_pdf() {
|
||||
|
||||
let cancel = Arc::new(AtomicBool::new(true)); // pre-set — abort immediately
|
||||
|
||||
let result = kebab_app::ingest_with_config_cancellable(
|
||||
let result = kebab_app::ingest_with_config(
|
||||
env.config.clone(),
|
||||
env.scope(),
|
||||
false,
|
||||
None,
|
||||
Some(cancel),
|
||||
kebab_app::IngestOpts {
|
||||
cancel: Some(cancel),
|
||||
..Default::default()
|
||||
},
|
||||
);
|
||||
// Both Ok (pre-cancel exit) and Err (eager OCR engine fail) are acceptable —
|
||||
// key assertion is no panic/deadlock.
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
//! Integration coverage for `ingest_with_config_progress` —
|
||||
//! Integration coverage for streaming ingest progress via `IngestOpts` —
|
||||
//! exercises the streaming progress channel against the same lexical
|
||||
//! fixture used by `ingest_lexical.rs`.
|
||||
|
||||
@@ -13,14 +13,19 @@ use kebab_core::IngestItemKind;
|
||||
fn run_with_progress() -> Vec<IngestEvent> {
|
||||
let env = TestEnv::lexical_only();
|
||||
let (tx, rx) = mpsc::channel::<IngestEvent>();
|
||||
let report =
|
||||
kebab_app::ingest_with_config_progress(env.config.clone(), env.scope(), false, Some(tx))
|
||||
.unwrap();
|
||||
let report = kebab_app::ingest_with_config(
|
||||
env.config.clone(),
|
||||
env.scope(),
|
||||
kebab_app::IngestOpts {
|
||||
progress: Some(tx),
|
||||
..Default::default()
|
||||
},
|
||||
)
|
||||
.unwrap();
|
||||
assert_eq!(report.scanned, 3);
|
||||
assert_eq!(report.new, 3);
|
||||
|
||||
// Drain until the sender (held inside `ingest_with_config_progress`)
|
||||
// is dropped on return.
|
||||
// Drain until the sender (held inside ingest_with_config) is dropped on return.
|
||||
let mut events = Vec::new();
|
||||
while let Ok(ev) = rx.recv() {
|
||||
events.push(ev);
|
||||
@@ -142,13 +147,18 @@ fn progress_event_sequence_matches_design_section_2_4a() {
|
||||
|
||||
#[test]
|
||||
fn ingest_with_config_progress_none_matches_ingest_with_config() {
|
||||
// Forwarding wrapper: `ingest_with_config(...)` and
|
||||
// `ingest_with_config_progress(..., None)` must produce identical
|
||||
// reports modulo wall-clock duration.
|
||||
// `ingest_with_config(...)` with no progress must produce identical
|
||||
// reports to a call with progress=None — modulo wall-clock duration.
|
||||
let env = TestEnv::lexical_only();
|
||||
let r_none =
|
||||
kebab_app::ingest_with_config_progress(env.config.clone(), env.scope(), true, None)
|
||||
.unwrap();
|
||||
let r_none = kebab_app::ingest_with_config(
|
||||
env.config.clone(),
|
||||
env.scope(),
|
||||
kebab_app::IngestOpts {
|
||||
summary_only: true,
|
||||
..Default::default()
|
||||
},
|
||||
)
|
||||
.unwrap();
|
||||
assert_eq!(r_none.scanned, 3);
|
||||
assert_eq!(r_none.new, 3);
|
||||
}
|
||||
@@ -160,9 +170,16 @@ fn dropped_receiver_does_not_panic_or_fail_ingest() {
|
||||
let env = TestEnv::lexical_only();
|
||||
let (tx, rx) = mpsc::channel::<IngestEvent>();
|
||||
drop(rx);
|
||||
let report =
|
||||
kebab_app::ingest_with_config_progress(env.config.clone(), env.scope(), true, Some(tx))
|
||||
.unwrap();
|
||||
let report = kebab_app::ingest_with_config(
|
||||
env.config.clone(),
|
||||
env.scope(),
|
||||
kebab_app::IngestOpts {
|
||||
progress: Some(tx),
|
||||
summary_only: true,
|
||||
..Default::default()
|
||||
},
|
||||
)
|
||||
.unwrap();
|
||||
assert_eq!(report.scanned, 3);
|
||||
}
|
||||
|
||||
@@ -172,7 +189,7 @@ fn dropped_receiver_does_not_panic_or_fail_ingest() {
|
||||
/// Manual invoke:
|
||||
/// ```
|
||||
/// KEBAB_PDF_OCR_ENABLED=true \
|
||||
/// KEBAB_PDF_OCR_ENDPOINT=http://192.168.0.47:11434 \
|
||||
/// KEBAB_OCR_ENDPOINT=http://192.168.0.47:11434 \
|
||||
/// cargo test -p kebab-app --test ingest_progress \
|
||||
/// --ignored pdf_ocr_progress_emits_started_finished_events
|
||||
/// ```
|
||||
@@ -197,7 +214,8 @@ fn pdf_ocr_progress_emits_started_finished_events() {
|
||||
config.models.embedding.provider = "none".to_string();
|
||||
config.models.embedding.dimensions = 0;
|
||||
config.ingest.pdf.ocr.enabled = true;
|
||||
if let Ok(endpoint) = std::env::var("KEBAB_PDF_OCR_ENDPOINT") {
|
||||
// v5: shared KEBAB_OCR_* env (manual harness reads it directly).
|
||||
if let Ok(endpoint) = std::env::var("KEBAB_OCR_ENDPOINT") {
|
||||
config.ingest.pdf.ocr.endpoint = Some(endpoint);
|
||||
}
|
||||
|
||||
@@ -207,8 +225,15 @@ fn pdf_ocr_progress_emits_started_finished_events() {
|
||||
};
|
||||
|
||||
let (tx, rx) = mpsc::channel::<IngestEvent>();
|
||||
let _report = kebab_app::ingest_with_config_progress(config, scope, false, Some(tx))
|
||||
.expect("ingest_with_config_progress");
|
||||
let _report = kebab_app::ingest_with_config(
|
||||
config,
|
||||
scope,
|
||||
kebab_app::IngestOpts {
|
||||
progress: Some(tx),
|
||||
..Default::default()
|
||||
},
|
||||
)
|
||||
.expect("ingest_with_config");
|
||||
|
||||
let events: Vec<_> = rx.iter().collect();
|
||||
|
||||
|
||||
@@ -52,6 +52,8 @@ fn extract_and_ocr(
|
||||
asset: &asset,
|
||||
workspace_root,
|
||||
config: &config,
|
||||
source_id: None,
|
||||
source_trust: None,
|
||||
};
|
||||
let mut canonical = PdfTextExtractor::new().extract(&ctx, bytes).unwrap();
|
||||
let opts = PdfOcrOpts {
|
||||
|
||||
@@ -72,7 +72,7 @@ fn seed_ocr_events(env: &TestEnv, store: &SqliteStore) {
|
||||
|
||||
fn open_app_with_seeded_events(env: &TestEnv) -> App {
|
||||
let app = env.app();
|
||||
let store = SqliteStore::open(&env.config).expect("open store for seed");
|
||||
let store = SqliteStore::open(&env.config.storage).expect("open store for seed");
|
||||
store.run_migrations().expect("run migrations for seed");
|
||||
seed_ocr_events(env, &store);
|
||||
app
|
||||
|
||||
@@ -49,6 +49,8 @@ fn extract_canonical_from_bytes(bytes: &[u8]) -> CanonicalDocument {
|
||||
asset: &asset,
|
||||
workspace_root,
|
||||
config: &config,
|
||||
source_id: None,
|
||||
source_trust: None,
|
||||
};
|
||||
PdfTextExtractor::new().extract(&ctx, bytes).unwrap()
|
||||
}
|
||||
|
||||
@@ -66,7 +66,7 @@ async fn ingest_dual_write_doc_id_matches_ndjson() {
|
||||
std::fs::copy(scanned_pdf_src(), &dest).expect("copy scanned PDF");
|
||||
|
||||
// Run ingest
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), false).expect("ingest");
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default()).expect("ingest");
|
||||
|
||||
// Read ndjson log
|
||||
let log_files: Vec<_> = std::fs::read_dir(&log_dir)
|
||||
|
||||
@@ -141,7 +141,7 @@ fn ingest_3_page_pdf_produces_one_doc_and_per_page_chunks() {
|
||||
write_pdf(&env.workspace_root, "three.pdf", &bytes);
|
||||
let cfg = cfg_with_pdf(&env);
|
||||
|
||||
let report = kebab_app::ingest_with_config(cfg.clone(), env.scope(), false)
|
||||
let report = kebab_app::ingest_with_config(cfg.clone(), env.scope(), kebab_app::IngestOpts::default())
|
||||
.expect("PDF ingest must succeed");
|
||||
|
||||
assert_eq!(report.errors, 0);
|
||||
@@ -203,7 +203,7 @@ fn re_ingest_identical_pdf_produces_unchanged_with_same_doc_id() {
|
||||
write_pdf(&env.workspace_root, "stable.pdf", &bytes);
|
||||
let cfg = cfg_with_pdf(&env);
|
||||
|
||||
let report1 = kebab_app::ingest_with_config(cfg.clone(), env.scope(), false).unwrap();
|
||||
let report1 = kebab_app::ingest_with_config(cfg.clone(), env.scope(), kebab_app::IngestOpts::default()).unwrap();
|
||||
let item1 = report1
|
||||
.items
|
||||
.as_ref()
|
||||
@@ -214,7 +214,7 @@ fn re_ingest_identical_pdf_produces_unchanged_with_same_doc_id() {
|
||||
.unwrap();
|
||||
assert_eq!(item1.kind, IngestItemKind::New);
|
||||
|
||||
let report2 = kebab_app::ingest_with_config(cfg.clone(), env.scope(), false).unwrap();
|
||||
let report2 = kebab_app::ingest_with_config(cfg.clone(), env.scope(), kebab_app::IngestOpts::default()).unwrap();
|
||||
let item2 = report2
|
||||
.items
|
||||
.unwrap()
|
||||
@@ -238,7 +238,7 @@ fn re_ingest_edited_pdf_produces_new_doc_id() {
|
||||
std::fs::write(&path, &bytes_v1).unwrap();
|
||||
let cfg = cfg_with_pdf(&env);
|
||||
|
||||
let report_v1 = kebab_app::ingest_with_config(cfg.clone(), env.scope(), false).unwrap();
|
||||
let report_v1 = kebab_app::ingest_with_config(cfg.clone(), env.scope(), kebab_app::IngestOpts::default()).unwrap();
|
||||
let id_v1 = report_v1
|
||||
.items
|
||||
.as_ref()
|
||||
@@ -253,7 +253,7 @@ fn re_ingest_edited_pdf_produces_new_doc_id() {
|
||||
let bytes_v2 = build_text_pdf(&[Some("VERSION TWO entirely different body content.")]);
|
||||
std::fs::write(&path, &bytes_v2).unwrap();
|
||||
|
||||
let report_v2 = kebab_app::ingest_with_config(cfg.clone(), env.scope(), false).unwrap();
|
||||
let report_v2 = kebab_app::ingest_with_config(cfg.clone(), env.scope(), kebab_app::IngestOpts::default()).unwrap();
|
||||
let item_v2 = report_v2
|
||||
.items
|
||||
.as_ref()
|
||||
@@ -278,7 +278,7 @@ fn encrypted_pdf_fails_with_qpdf_hint() {
|
||||
write_pdf(&env.workspace_root, "secret.pdf", &bytes);
|
||||
let cfg = cfg_with_pdf(&env);
|
||||
|
||||
let report = kebab_app::ingest_with_config(cfg, env.scope(), false).unwrap();
|
||||
let report = kebab_app::ingest_with_config(cfg, env.scope(), kebab_app::IngestOpts::default()).unwrap();
|
||||
assert_eq!(
|
||||
report.errors, 1,
|
||||
"encrypted PDF must increment errors exactly once"
|
||||
@@ -308,7 +308,7 @@ fn corrupt_pdf_fails_without_storing() {
|
||||
write_pdf(&env.workspace_root, "corrupt.pdf", &bytes);
|
||||
let cfg = cfg_with_pdf(&env);
|
||||
|
||||
let report = kebab_app::ingest_with_config(cfg.clone(), env.scope(), false).unwrap();
|
||||
let report = kebab_app::ingest_with_config(cfg.clone(), env.scope(), kebab_app::IngestOpts::default()).unwrap();
|
||||
assert_eq!(
|
||||
report.errors, 1,
|
||||
"corrupt PDF must increment errors exactly once"
|
||||
@@ -342,7 +342,7 @@ fn mixed_page_pdf_stores_asset_with_scanned_candidate_warning() {
|
||||
write_pdf(&env.workspace_root, "mixed.pdf", &bytes);
|
||||
let cfg = cfg_with_pdf(&env);
|
||||
|
||||
let report = kebab_app::ingest_with_config(cfg.clone(), env.scope(), false).unwrap();
|
||||
let report = kebab_app::ingest_with_config(cfg.clone(), env.scope(), kebab_app::IngestOpts::default()).unwrap();
|
||||
assert_eq!(
|
||||
report.errors, 0,
|
||||
"scanned candidate is a Warning, not Error"
|
||||
@@ -413,7 +413,7 @@ fn ingest_report_arithmetic_invariant_holds_with_corrupt_pdf() {
|
||||
write_pdf(&env.workspace_root, "broken.pdf", &corrupt_pdf());
|
||||
let cfg = cfg_with_pdf(&env);
|
||||
|
||||
let report = kebab_app::ingest_with_config(cfg, env.scope(), false).unwrap();
|
||||
let report = kebab_app::ingest_with_config(cfg, env.scope(), kebab_app::IngestOpts::default()).unwrap();
|
||||
let total = report.new + report.updated + report.skipped + report.errors;
|
||||
assert_eq!(
|
||||
report.scanned, total,
|
||||
@@ -439,7 +439,7 @@ fn long_pdf_round_trips_through_lexical_pipeline() {
|
||||
write_pdf(&env.workspace_root, "long.pdf", &bytes);
|
||||
let cfg = cfg_with_pdf(&env);
|
||||
|
||||
let report = kebab_app::ingest_with_config(cfg.clone(), env.scope(), false).unwrap();
|
||||
let report = kebab_app::ingest_with_config(cfg.clone(), env.scope(), kebab_app::IngestOpts::default()).unwrap();
|
||||
assert_eq!(report.errors, 0);
|
||||
let pdf_item = report
|
||||
.items
|
||||
@@ -470,7 +470,7 @@ fn inspect_doc_surfaces_page_spans() {
|
||||
write_pdf(&env.workspace_root, "inspect.pdf", &bytes);
|
||||
let cfg = cfg_with_pdf(&env);
|
||||
|
||||
let report = kebab_app::ingest_with_config(cfg.clone(), env.scope(), false).unwrap();
|
||||
let report = kebab_app::ingest_with_config(cfg.clone(), env.scope(), kebab_app::IngestOpts::default()).unwrap();
|
||||
let pdf_item = report
|
||||
.items
|
||||
.as_ref()
|
||||
|
||||
@@ -16,14 +16,14 @@
|
||||
mod common;
|
||||
|
||||
use common::TestEnv;
|
||||
use kebab_app::IngestOpts;
|
||||
use kebab_app::{IngestOpts, ingest_with_config};
|
||||
use kebab_app::reset::{ResetScope, execute};
|
||||
use kebab_core::{DocFilter, DocumentStore, SourceScope};
|
||||
|
||||
/// Open the SqliteStore and list all `workspace_path` values.
|
||||
fn list_doc_paths(env: &TestEnv) -> Vec<String> {
|
||||
use kebab_store_sqlite::SqliteStore;
|
||||
let store = SqliteStore::open(&env.config).unwrap();
|
||||
let store = SqliteStore::open(&env.config.storage).unwrap();
|
||||
store.run_migrations().unwrap();
|
||||
store
|
||||
.list_documents(&DocFilter::default())
|
||||
@@ -51,13 +51,8 @@ fn reset_orphans_only_purges_out_of_scope_docs() {
|
||||
include: vec!["**/*.rs".to_string()],
|
||||
exclude: env.config.workspace.exclude.clone(),
|
||||
};
|
||||
let first = kebab_app::ingest_with_config_opts(
|
||||
env.config.clone(),
|
||||
wide_scope,
|
||||
false,
|
||||
IngestOpts::default(),
|
||||
)
|
||||
.expect("first ingest must succeed");
|
||||
let first = ingest_with_config(env.config.clone(), wide_scope, IngestOpts::default())
|
||||
.expect("first ingest must succeed");
|
||||
// The fixture workspace may contain other .rs files — just assert we
|
||||
// got at least 3 new docs (our a.rs, b.rs, c.rs).
|
||||
assert!(first.new >= 3, "expected at least 3 new docs: {first:?}");
|
||||
|
||||
@@ -32,7 +32,7 @@ fn schema_models_active_arrays_empty_on_empty_corpus() {
|
||||
std::fs::create_dir_all(&workspace).unwrap();
|
||||
let cfg = minimal_config(dir.path(), &workspace);
|
||||
|
||||
let store = kebab_store_sqlite::SqliteStore::open(&cfg).unwrap();
|
||||
let store = kebab_store_sqlite::SqliteStore::open(&cfg.storage).unwrap();
|
||||
store.run_migrations().unwrap();
|
||||
drop(store);
|
||||
|
||||
@@ -58,7 +58,7 @@ fn schema_emits_active_parsers_and_chunkers_array_after_ingest() {
|
||||
let cfg = minimal_config(dir.path(), &workspace);
|
||||
let scope = minimal_scope(&workspace);
|
||||
|
||||
kebab_app::ingest_with_config(cfg.clone(), scope, false).unwrap();
|
||||
kebab_app::ingest_with_config(cfg.clone(), scope, kebab_app::IngestOpts::default()).unwrap();
|
||||
|
||||
let s = schema_with_config(&cfg).unwrap();
|
||||
assert!(
|
||||
|
||||
@@ -40,7 +40,7 @@ fn schema_report_reflects_freshly_ingested_kb() {
|
||||
|
||||
let config = minimal_config(&data_dir, &workspace_root);
|
||||
let _report =
|
||||
kebab_app::ingest_with_config(config.clone(), minimal_scope(&workspace_root), false)
|
||||
kebab_app::ingest_with_config(config.clone(), minimal_scope(&workspace_root), kebab_app::IngestOpts::default())
|
||||
.unwrap();
|
||||
|
||||
let schema = kebab_app::schema_with_config(&config).unwrap();
|
||||
@@ -100,7 +100,7 @@ fn schema_report_on_empty_kb_has_zero_counts() {
|
||||
// Run ingest over the empty workspace — creates kebab.sqlite, runs
|
||||
// migrations, records 0 docs. schema_with_config can then open_existing.
|
||||
let report =
|
||||
kebab_app::ingest_with_config(config.clone(), minimal_scope(&workspace_root), false)
|
||||
kebab_app::ingest_with_config(config.clone(), minimal_scope(&workspace_root), kebab_app::IngestOpts::default())
|
||||
.unwrap();
|
||||
assert_eq!(report.new, 0, "empty workspace should yield 0 new docs");
|
||||
|
||||
|
||||
@@ -24,7 +24,7 @@ fn korean_lexical_query_returns_korean_document() {
|
||||
.expect("write Korean fixture doc");
|
||||
|
||||
// Ingest — lexical_only() disables fastembed so no AVX required.
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), true)
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts { summary_only: true, ..Default::default() })
|
||||
.expect("ingest must succeed");
|
||||
|
||||
// Lexical search for "러스트" — must return the Korean document.
|
||||
@@ -72,7 +72,7 @@ fn lexical_multi_token_korean_query_hits() {
|
||||
.join("hash-table.md");
|
||||
std::fs::copy(&src, &dest).expect("copy korean fixture");
|
||||
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), true)
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts { summary_only: true, ..Default::default() })
|
||||
.expect("ingest must succeed");
|
||||
|
||||
let hits =
|
||||
@@ -108,7 +108,7 @@ fn lexical_mixed_korean_english_multi_token_query_hits() {
|
||||
)
|
||||
.expect("write rust-hash fixture");
|
||||
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), true)
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts { summary_only: true, ..Default::default() })
|
||||
.expect("ingest must succeed");
|
||||
|
||||
let hits =
|
||||
@@ -144,7 +144,7 @@ fn korean_morphological_2char_query_lexical_mode() {
|
||||
)
|
||||
.expect("write korean-wiki fixture");
|
||||
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), true)
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts { summary_only: true, ..Default::default() })
|
||||
.expect("ingest must succeed");
|
||||
|
||||
let hits = kebab_app::search_with_config(env.config.clone(), common::lexical_query("한국"))
|
||||
@@ -176,7 +176,7 @@ fn korean_morphological_mixed_english_korean_query() {
|
||||
)
|
||||
.expect("write rust-optimization fixture");
|
||||
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), true)
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts { summary_only: true, ..Default::default() })
|
||||
.expect("ingest must succeed");
|
||||
|
||||
let hits = kebab_app::search_with_config(env.config.clone(), common::lexical_query("Rust"))
|
||||
|
||||
@@ -8,7 +8,7 @@ use common::TestEnv;
|
||||
#[test]
|
||||
fn lexical_search_returns_hits_after_ingest() {
|
||||
let env = TestEnv::lexical_only();
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), true).unwrap();
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts { summary_only: true, ..Default::default() }).unwrap();
|
||||
|
||||
// "Ownership" appears as a heading + paragraph in intro.md and
|
||||
// matches FTS5 default tokenizer easily.
|
||||
@@ -34,7 +34,7 @@ fn lexical_search_returns_hits_after_ingest() {
|
||||
#[test]
|
||||
fn lexical_search_empty_query_returns_empty() {
|
||||
let env = TestEnv::lexical_only();
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), true).unwrap();
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts { summary_only: true, ..Default::default() }).unwrap();
|
||||
let hits =
|
||||
kebab_app::search_with_config(env.config.clone(), common::lexical_query(" ")).unwrap();
|
||||
assert!(hits.is_empty(), "blank query must short-circuit empty");
|
||||
@@ -46,7 +46,7 @@ fn lexical_search_empty_query_returns_empty() {
|
||||
#[test]
|
||||
fn cached_search_returns_same_hits_on_repeat() {
|
||||
let env = TestEnv::lexical_only();
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), true).unwrap();
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts { summary_only: true, ..Default::default() }).unwrap();
|
||||
let app = kebab_app::App::open_with_config(env.config.clone()).unwrap();
|
||||
let first = app.search(common::lexical_query("ownership")).unwrap();
|
||||
assert!(!first.is_empty(), "first call must return ≥1 hit");
|
||||
@@ -68,7 +68,7 @@ fn cached_search_returns_same_hits_on_repeat() {
|
||||
#[test]
|
||||
fn cache_key_normalization_treats_case_and_whitespace_as_equivalent() {
|
||||
let env = TestEnv::lexical_only();
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), true).unwrap();
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts { summary_only: true, ..Default::default() }).unwrap();
|
||||
let app = kebab_app::App::open_with_config(env.config.clone()).unwrap();
|
||||
let plain = app.search(common::lexical_query("ownership")).unwrap();
|
||||
let upper = app.search(common::lexical_query("OWNERSHIP")).unwrap();
|
||||
@@ -86,7 +86,7 @@ fn cache_key_normalization_treats_case_and_whitespace_as_equivalent() {
|
||||
#[test]
|
||||
fn search_uncached_returns_same_hits_as_cached() {
|
||||
let env = TestEnv::lexical_only();
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), true).unwrap();
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts { summary_only: true, ..Default::default() }).unwrap();
|
||||
let cached =
|
||||
kebab_app::search_with_config(env.config.clone(), common::lexical_query("ownership"))
|
||||
.unwrap();
|
||||
@@ -107,7 +107,7 @@ fn search_uncached_returns_same_hits_as_cached() {
|
||||
#[test]
|
||||
fn first_ingest_bumps_corpus_revision() {
|
||||
let env = TestEnv::lexical_only();
|
||||
let store_before = kebab_store_sqlite::SqliteStore::open(&env.config).unwrap();
|
||||
let store_before = kebab_store_sqlite::SqliteStore::open(&env.config.storage).unwrap();
|
||||
store_before.run_migrations().unwrap();
|
||||
// V004 seeds 0; V009 + V010 + V011 migrations each bump by 1 to
|
||||
// invalidate stale LRU caches (spec §5.2). Baseline before ingest = 3.
|
||||
@@ -116,13 +116,13 @@ fn first_ingest_bumps_corpus_revision() {
|
||||
let baseline = store_before.corpus_revision();
|
||||
assert_eq!(baseline, 3, "fresh store post-V011 baseline = 3");
|
||||
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), true).unwrap();
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts { summary_only: true, ..Default::default() }).unwrap();
|
||||
assert!(
|
||||
report.new + report.updated > 0,
|
||||
"first ingest must commit ≥1 doc"
|
||||
);
|
||||
|
||||
let store_after = kebab_store_sqlite::SqliteStore::open(&env.config).unwrap();
|
||||
let store_after = kebab_store_sqlite::SqliteStore::open(&env.config.storage).unwrap();
|
||||
assert!(
|
||||
store_after.corpus_revision() > baseline,
|
||||
"ingest commit must bump corpus_revision past baseline {baseline} (got {})",
|
||||
@@ -133,7 +133,7 @@ fn first_ingest_bumps_corpus_revision() {
|
||||
#[test]
|
||||
fn vector_mode_with_provider_none_errors_clearly() {
|
||||
let env = TestEnv::lexical_only();
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), true).unwrap();
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts { summary_only: true, ..Default::default() }).unwrap();
|
||||
|
||||
let q = kebab_core::SearchQuery {
|
||||
text: "ownership".to_string(),
|
||||
|
||||
@@ -21,7 +21,7 @@ fn lexical_query_owner() -> kebab_core::SearchQuery {
|
||||
#[test]
|
||||
fn fresh_doc_is_not_stale_with_default_threshold() {
|
||||
let env = TestEnv::lexical_only();
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), true).unwrap();
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts { summary_only: true, ..Default::default() }).unwrap();
|
||||
|
||||
let app = kebab_app::App::open_with_config(env.config.clone()).unwrap();
|
||||
let hits = app.search(lexical_query_owner()).unwrap();
|
||||
@@ -43,7 +43,7 @@ fn threshold_zero_disables_staleness() {
|
||||
let mut env = TestEnv::lexical_only();
|
||||
env.config.search.stale_threshold_days = 0;
|
||||
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), true).unwrap();
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts { summary_only: true, ..Default::default() }).unwrap();
|
||||
common::backdate_document_updated_at(&env, "intro.md", 365);
|
||||
|
||||
let app = kebab_app::App::open_with_config(env.config.clone()).unwrap();
|
||||
@@ -66,7 +66,7 @@ fn old_doc_marked_stale() {
|
||||
let mut env = TestEnv::lexical_only();
|
||||
env.config.search.stale_threshold_days = 30;
|
||||
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), true).unwrap();
|
||||
kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts { summary_only: true, ..Default::default() }).unwrap();
|
||||
common::backdate_document_updated_at(&env, "intro.md", 60);
|
||||
|
||||
let app = kebab_app::App::open_with_config(env.config.clone()).unwrap();
|
||||
|
||||
@@ -29,7 +29,7 @@ fn ingest_then_hybrid_search_returns_hits() {
|
||||
require_avx_or_panic();
|
||||
|
||||
let env = TestEnv::with_embeddings();
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), true).unwrap();
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts { summary_only: true, ..Default::default() }).unwrap();
|
||||
assert_eq!(report.errors, 0, "no per-file errors: {report:?}");
|
||||
assert_eq!(report.new, 3);
|
||||
|
||||
@@ -55,7 +55,7 @@ fn ingest_then_vector_search_carries_embedding_model() {
|
||||
require_avx_or_panic();
|
||||
|
||||
let env = TestEnv::with_embeddings();
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), true).unwrap();
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts { summary_only: true, ..Default::default() }).unwrap();
|
||||
assert_eq!(report.errors, 0, "no per-file errors: {report:?}");
|
||||
assert_eq!(report.new, 3);
|
||||
|
||||
|
||||
@@ -13,7 +13,7 @@ fn unsupported_extension_skip_carries_warning_and_is_aggregated() {
|
||||
std::fs::write(workspace_root.join("legacy.docx"), b"unsupported").unwrap();
|
||||
std::fs::write(workspace_root.join("Makefile"), b"unsupported").unwrap();
|
||||
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), false).unwrap();
|
||||
let report = kebab_app::ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default()).unwrap();
|
||||
|
||||
let items = report.items.as_ref().expect("items array populated");
|
||||
let docx_item = items
|
||||
|
||||
@@ -45,7 +45,7 @@ fn twin_files_fetch_span_uses_correct_asset() {
|
||||
|
||||
// Ingest all files (fixture workspace + our two new twins).
|
||||
let report =
|
||||
ingest_with_config(env.config.clone(), env.scope(), false).expect("ingest must succeed");
|
||||
ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default()).expect("ingest must succeed");
|
||||
assert_eq!(report.errors, 0, "no ingest errors; report={report:?}");
|
||||
|
||||
// Both twin paths must appear as New in the report.
|
||||
@@ -72,7 +72,7 @@ fn twin_files_fetch_span_uses_correct_asset() {
|
||||
// Resolve doc_ids for both workspace paths.
|
||||
// The ingest layer normalises workspace_path to the path relative to
|
||||
// workspace_root (e.g. "src_a/note.md"), so we look up by that form.
|
||||
let store = kebab_store_sqlite::SqliteStore::open(&env.config).unwrap();
|
||||
let store = kebab_store_sqlite::SqliteStore::open(&env.config.storage).unwrap();
|
||||
store.run_migrations().unwrap();
|
||||
|
||||
// Find the twin items by matching on suffix so the test is robust to
|
||||
@@ -146,7 +146,7 @@ fn twin_files_fetch_span_uses_correct_asset() {
|
||||
// re-check. Pre-fix this was the scenario that triggered the bug:
|
||||
// after the second ingest the asset row's workspace_path could point
|
||||
// at either twin, making one twin's span fetch behave incorrectly.
|
||||
let report2 = ingest_with_config(env.config.clone(), env.scope(), false)
|
||||
let report2 = ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default())
|
||||
.expect("second ingest must succeed");
|
||||
assert_eq!(
|
||||
report2.errors, 0,
|
||||
|
||||
@@ -36,7 +36,7 @@ fn twin_files_second_ingest_is_unchanged() {
|
||||
std::fs::write(pkg_b.join("__init__.py"), content).unwrap();
|
||||
|
||||
// First ingest — both files must be New.
|
||||
let first = ingest_with_config(env.config.clone(), env.scope(), false)
|
||||
let first = ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default())
|
||||
.expect("first ingest must succeed");
|
||||
assert_eq!(first.errors, 0, "first ingest: no errors; report={first:?}");
|
||||
|
||||
@@ -59,7 +59,7 @@ fn twin_files_second_ingest_is_unchanged() {
|
||||
}
|
||||
|
||||
// Second ingest — same files, same content → both must be Unchanged.
|
||||
let second = ingest_with_config(env.config.clone(), env.scope(), false)
|
||||
let second = ingest_with_config(env.config.clone(), env.scope(), kebab_app::IngestOpts::default())
|
||||
.expect("second ingest must succeed");
|
||||
assert_eq!(
|
||||
second.errors, 0,
|
||||
|
||||
@@ -171,6 +171,8 @@ fn extract_cpp_fixture() -> CanonicalDocument {
|
||||
asset: &asset,
|
||||
workspace_root: &root,
|
||||
config: &cfg,
|
||||
source_id: None,
|
||||
source_trust: None,
|
||||
};
|
||||
CppAstExtractor::new()
|
||||
.extract(&ctx, src.as_bytes())
|
||||
|
||||
@@ -23,10 +23,6 @@ kebab-app = { path = "../kebab-app" }
|
||||
# kb-cli → kb-eval directly; documented in
|
||||
# `tasks/p5/p5-2-metrics-compare.md`.
|
||||
kebab-eval = { path = "../kebab-eval" }
|
||||
# P9-1: Ratatui shell. UI consumes `kebab-app` only — `kebab-tui`
|
||||
# enforces the §8 boundary in its own Cargo.toml; kb-cli just
|
||||
# launches it.
|
||||
kebab-tui = { path = "../kebab-tui" }
|
||||
# p9-fb-30: MCP stdio server. `Cmd::Mcp` delegates entirely to this crate.
|
||||
kebab-mcp = { path = "../kebab-mcp" }
|
||||
anyhow = { workspace = true }
|
||||
@@ -51,10 +47,5 @@ tempfile = { workspace = true }
|
||||
rusqlite = { workspace = true }
|
||||
time = { workspace = true }
|
||||
|
||||
[features]
|
||||
# opt-in (macOS): build the `kebab` binary with candle on the Apple Silicon GPU.
|
||||
# cargo build --release --features embed_metal
|
||||
embed_metal = ["kebab-app/embed_metal"]
|
||||
|
||||
[lints]
|
||||
workspace = true
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
//! `kebab ingest` SIGINT (Ctrl-C) handler — flips a shared
|
||||
//! `Arc<AtomicBool>` so `kebab_app::ingest_with_config_cancellable`
|
||||
//! can break at the next step boundary.
|
||||
//! `Arc<AtomicBool>` so `kebab_app::ingest_with_config`
|
||||
//! can break at the next step boundary (via `IngestOpts::cancel`).
|
||||
//!
|
||||
//! Per spec §10: the second Ctrl-C is a hard exit (130 = SIGINT
|
||||
//! conventional). We count signal arrivals via a private atomic and
|
||||
@@ -25,8 +25,8 @@ use std::sync::atomic::{AtomicBool, AtomicU8, Ordering};
|
||||
|
||||
/// Install a SIGINT handler that:
|
||||
/// - on first signal: sets `cancel.store(true)` so the cooperative
|
||||
/// cancel loop in `kebab_app::ingest_with_config_cancellable`
|
||||
/// breaks at its next step boundary.
|
||||
/// cancel loop in `kebab_app::ingest_with_config` (via
|
||||
/// `IngestOpts::cancel`) breaks at its next step boundary.
|
||||
/// - on second signal: hard-exits with code 130 (SIGINT
|
||||
/// convention).
|
||||
///
|
||||
|
||||
@@ -271,16 +271,6 @@ enum Cmd {
|
||||
#[arg(long)]
|
||||
hide_citations: bool,
|
||||
|
||||
/// p9-fb-18: persistent multi-turn chat session id. First call
|
||||
/// auto-creates the session in SQLite (`chat_sessions`), each
|
||||
/// subsequent call with the same id loads prior turns as
|
||||
/// history and appends the new Q/A. Without this flag, ask
|
||||
/// is single-shot (no persistence). The session id is
|
||||
/// caller-supplied — pick anything stable per conversation
|
||||
/// (e.g. `kebab-rust-async-2026-05`).
|
||||
#[arg(long, value_name = "ID")]
|
||||
session: Option<String>,
|
||||
|
||||
/// p9-fb-33: emit ndjson `answer_event.v1` events on stderr
|
||||
/// while streaming. Final stdout line is the existing
|
||||
/// `answer.v1`. Off by default to preserve final-only behavior.
|
||||
@@ -343,10 +333,6 @@ enum Cmd {
|
||||
/// Print introspection report (wire schemas, capabilities, model versions, stats).
|
||||
Schema,
|
||||
|
||||
/// Launch the Ratatui shell (P9-1 — Library pane only; search /
|
||||
/// ask / inspect panes land with p9-2 / p9-3 / p9-4).
|
||||
Tui,
|
||||
|
||||
/// Eval suite (placeholder; lands in P9).
|
||||
Eval {
|
||||
#[command(subcommand)]
|
||||
@@ -664,14 +650,9 @@ fn run(cli: &Cli) -> anyhow::Result<()> {
|
||||
let mode = progress::ProgressMode::from_flags(cli.json, cli.quiet, plain_env);
|
||||
|
||||
// Surface the active embedding backend/device on the terminal so the
|
||||
// user sees it without grepping kb.log (the per-device tracing line
|
||||
// only lands in the log file at --verbose). Suppressed under
|
||||
// --json/--quiet. The Metal note reflects the build (`embed_metal`);
|
||||
// the confirmed runtime device is in kb.log (`candle device = ...`).
|
||||
// user sees it without grepping kb.log. Suppressed under --json/--quiet.
|
||||
if !cli.json && !cli.quiet {
|
||||
let backend = match cfg.models.embedding.provider.as_str() {
|
||||
"candle" if cfg!(feature = "embed_metal") => "candle (Metal/GPU 빌드)",
|
||||
"candle" => "candle (CPU, 순수 Rust)",
|
||||
"fastembed" | "onnx" | "" => "fastembed (onnxruntime)",
|
||||
"none" => "비활성 (lexical-only)",
|
||||
other => other,
|
||||
@@ -691,14 +672,14 @@ fn run(cli: &Cli) -> anyhow::Result<()> {
|
||||
|
||||
// p9-fb-23: use IngestOpts so force_reingest threads through
|
||||
// without churning the positional-arg list.
|
||||
let ingest_result = kebab_app::ingest_with_config_opts(
|
||||
let ingest_result = kebab_app::ingest_with_config(
|
||||
cfg,
|
||||
scope,
|
||||
*summary_only,
|
||||
kebab_app::IngestOpts {
|
||||
progress: Some(tx),
|
||||
cancel: Some(cancel_token),
|
||||
force_reingest: *force_reingest,
|
||||
summary_only: *summary_only,
|
||||
},
|
||||
);
|
||||
|
||||
@@ -845,7 +826,7 @@ fn run(cli: &Cli) -> anyhow::Result<()> {
|
||||
k,
|
||||
mode,
|
||||
explain: _,
|
||||
no_cache,
|
||||
no_cache: _,
|
||||
max_tokens,
|
||||
snippet_chars,
|
||||
cursor,
|
||||
@@ -1015,12 +996,8 @@ fn run(cli: &Cli) -> anyhow::Result<()> {
|
||||
cursor: cursor.clone(),
|
||||
trace: *trace,
|
||||
};
|
||||
// p9-fb-34: budget-aware path. --no-cache still bypasses the
|
||||
// App-level LRU; wire wrapper applies regardless.
|
||||
// p9-fb-34: budget-aware path.
|
||||
let app = kebab_app::App::open_with_config(cfg)?;
|
||||
if *no_cache {
|
||||
app.clear_search_cache();
|
||||
}
|
||||
let resp = app.search_with_opts(q, opts)?;
|
||||
|
||||
if cli.json {
|
||||
@@ -1123,7 +1100,6 @@ fn run(cli: &Cli) -> anyhow::Result<()> {
|
||||
seed,
|
||||
show_citations,
|
||||
hide_citations,
|
||||
session,
|
||||
stream,
|
||||
multi_hop,
|
||||
} => {
|
||||
@@ -1160,19 +1136,12 @@ fn run(cli: &Cli) -> anyhow::Result<()> {
|
||||
temperature: *temperature,
|
||||
seed: *seed,
|
||||
stream_sink: Some(tx),
|
||||
history: Vec::new(),
|
||||
conversation_id: None,
|
||||
turn_index: None,
|
||||
multi_hop: *multi_hop,
|
||||
};
|
||||
let cfg2 = cfg.clone();
|
||||
let q = query.clone();
|
||||
let session2 = session.clone();
|
||||
let handle = std::thread::spawn(move || -> anyhow::Result<kebab_core::Answer> {
|
||||
match session2.as_deref() {
|
||||
Some(sid) => kebab_app::ask_with_session_with_config(cfg2, sid, &q, opts),
|
||||
None => kebab_app::ask_with_config(cfg2, &q, opts),
|
||||
}
|
||||
kebab_app::ask_with_config(cfg2, &q, opts)
|
||||
});
|
||||
|
||||
// Drain receiver, write ndjson to stderr until
|
||||
@@ -1227,20 +1196,9 @@ fn run(cli: &Cli) -> anyhow::Result<()> {
|
||||
// takes the branch above; the TUI ask pane (P9-3)
|
||||
// wires up its own `mpsc::Sender`.
|
||||
stream_sink: None,
|
||||
// p9-fb-18: when `--session` is set, the facade
|
||||
// (`ask_with_session_with_config`) loads prior turns
|
||||
// from SQLite and stuffs them into AskOpts.history
|
||||
// before calling `ask_with_history`. Single-shot path
|
||||
// (no `--session`) keeps the empty defaults.
|
||||
history: Vec::new(),
|
||||
conversation_id: None,
|
||||
turn_index: None,
|
||||
multi_hop: *multi_hop,
|
||||
};
|
||||
let ans = match session.as_deref() {
|
||||
Some(sid) => kebab_app::ask_with_session_with_config(cfg, sid, query, opts)?,
|
||||
None => kebab_app::ask_with_config(cfg, query, opts)?,
|
||||
};
|
||||
let ans = kebab_app::ask_with_config(cfg, query, opts)?;
|
||||
if cli.json {
|
||||
println!("{}", serde_json::to_string(&wire::wire_answer(&ans))?);
|
||||
} else {
|
||||
@@ -1438,17 +1396,6 @@ fn run(cli: &Cli) -> anyhow::Result<()> {
|
||||
Ok(())
|
||||
}
|
||||
|
||||
Cmd::Tui => {
|
||||
// P9-1: Ratatui shell with Library pane. Search / Ask /
|
||||
// Inspect panes land in p9-2 / p9-3 / p9-4.
|
||||
let config = match cli.config.as_deref() {
|
||||
Some(path) => kebab_config::Config::load(Some(path))?,
|
||||
None => kebab_config::Config::load(None)?,
|
||||
};
|
||||
let mut app = kebab_tui::App::new(config)?;
|
||||
app.run()
|
||||
}
|
||||
|
||||
Cmd::Eval { what } => {
|
||||
let cfg = kebab_config::Config::load(cli.config.as_deref())?;
|
||||
match what {
|
||||
@@ -1860,8 +1807,6 @@ mod tests {
|
||||
latency_ms: 0,
|
||||
},
|
||||
created_at: OffsetDateTime::now_utc(),
|
||||
conversation_id: None,
|
||||
turn_index: None,
|
||||
hops: None,
|
||||
verification: None,
|
||||
}
|
||||
|
||||
@@ -16,7 +16,7 @@
|
||||
//! Each subprocess of the binary creates one `ProgressDisplay` and
|
||||
//! drives it from a background thread that drains an
|
||||
//! `mpsc::Receiver<IngestEvent>`. The thread terminates when the
|
||||
//! `Sender` end is dropped (i.e. when `ingest_with_config_progress`
|
||||
//! `Sender` end is dropped (i.e. when `ingest_with_config`
|
||||
//! returns).
|
||||
|
||||
use std::collections::HashMap;
|
||||
|
||||
@@ -337,8 +337,8 @@ mod tests {
|
||||
json_mode: true,
|
||||
ingest_progress: true,
|
||||
ingest_cancellation: true,
|
||||
rag_multi_turn: true,
|
||||
search_cache: true,
|
||||
rag_multi_turn: false,
|
||||
search_cache: false,
|
||||
incremental_ingest: true,
|
||||
streaming_ask: false,
|
||||
http_daemon: false,
|
||||
|
||||
@@ -64,7 +64,7 @@ rrf_k = 60
|
||||
snippet_chars = 220
|
||||
|
||||
[rag]
|
||||
prompt_template_version = "rag-v1"
|
||||
prompt_template_version = "rag-v4"
|
||||
score_gate = 0.30
|
||||
explain_default = false
|
||||
max_context_tokens = 8000
|
||||
|
||||
@@ -65,7 +65,7 @@ rrf_k = 60
|
||||
snippet_chars = 220
|
||||
|
||||
[rag]
|
||||
prompt_template_version = "rag-v1"
|
||||
prompt_template_version = "rag-v4"
|
||||
score_gate = 0.30
|
||||
explain_default = false
|
||||
max_context_tokens = 8000
|
||||
|
||||
@@ -90,7 +90,7 @@ snippet_chars = 220
|
||||
stale_threshold_days = {stale_threshold_days}
|
||||
|
||||
[rag]
|
||||
prompt_template_version = "rag-v1"
|
||||
prompt_template_version = "rag-v4"
|
||||
score_gate = 0.30
|
||||
explain_default = false
|
||||
max_context_tokens = 8000
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -9,7 +9,7 @@ use toml_edit::{DocumentMut, Item};
|
||||
|
||||
/// 현재 바이너리가 이해하는 config 스키마 버전. 마이그레이션 완료 시
|
||||
/// 사용자 파일의 `schema_version` 을 이 값으로 stamp 한다.
|
||||
pub const CURRENT_SCHEMA_VERSION: u32 = 4;
|
||||
pub const CURRENT_SCHEMA_VERSION: u32 = 5;
|
||||
|
||||
/// 한 번의 마이그레이션에서 발생한 개별 변경.
|
||||
#[derive(Clone, Debug, PartialEq, serde::Serialize)]
|
||||
@@ -80,6 +80,7 @@ fn section_comment(path: &str) -> Option<&'static str> {
|
||||
"rag" => "# 답변 생성: prompt 템플릿·score gate·NLI.",
|
||||
"ui" => "# TUI 팔레트·role 스타일.",
|
||||
"ingest" => "# 모든 형식 ingest 우산: 병렬도 + chunking/code/image/pdf.",
|
||||
"ingest.ocr" => "# 공유 OCR 엔진 설정(image/pdf 공통). 각 미디어 블록이 override.",
|
||||
"ingest.chunking" => "# 청크 크기·오버랩·heading 존중(전 형식 공통).",
|
||||
"ingest.code" => "# code ingest skip 정책(.gitignore 자동 honor).",
|
||||
"ingest.image" => "# 이미지 OCR + 캡션(기본 off, asset 당 모델 호출 비용).",
|
||||
@@ -99,9 +100,9 @@ fn key_comment(path: &str) -> Option<&'static str> {
|
||||
"workspace.root" => "색인 루트. 절대/~/${VAR}/상대(=이 파일 기준).",
|
||||
"workspace.exclude" => "denylist glob.",
|
||||
"storage.copy_threshold_mb" => "이 크기(MB) 초과 파일은 사본 대신 참조.",
|
||||
"models.embedding.provider" => "fastembed | candle | ollama | none.",
|
||||
"models.embedding.provider" => "fastembed | ollama | none.",
|
||||
"models.embedding.dimensions" => "모델 출력 차원. 틀리면 검색 0건.",
|
||||
"models.embedding.num_threads" => "candle 전용 CPU 스레드 cap(0=auto).",
|
||||
"models.embedding.num_threads" => "레거시 필드 (deprecated, 무시됨).",
|
||||
"models.embedding.endpoint" => "ollama provider 시 HTTP. 비우면 llm.endpoint fallback.",
|
||||
"models.llm.request_timeout_secs" => "단일 HTTP 상한. 0=즉시실패(비활성화 아님).",
|
||||
"ingest.max_parallel_extractors" => "동시 extractor 수.",
|
||||
@@ -142,6 +143,16 @@ pub fn annotated_default_document() -> DocumentMut {
|
||||
let pretty = toml::to_string_pretty(&defaults).expect("defaults serialize");
|
||||
let mut doc: DocumentMut = pretty.parse().expect("defaults parse as toml_edit");
|
||||
|
||||
// v5: 직렬화된 defaults 는 `[ingest.image.ocr]` 에 12개 엔진 키를 그대로
|
||||
// 담고 `[ingest.ocr]` 은 비어 있다(`SharedOcrEngineCfg::default()` = 전부
|
||||
// None). 참조 문서를 v5 canonical 형상으로 맞추기 위해 같은 통합을 적용한다
|
||||
// — 그러지 않으면 reconcile 이 마이그레이션으로 끌어올린 키를 image 블록에
|
||||
// 다시 추가해 통합을 되돌린다(그리고 image 의 effective engine 을 default 로
|
||||
// 덮어써 동작을 바꾼다). pdf 블록은 그대로 둔다(자기 default 가 image 와
|
||||
// 달라 override 로 유지).
|
||||
let mut discard = Vec::new();
|
||||
step_4_to_5(&mut doc, &mut discard);
|
||||
|
||||
// 헤더: 첫 최상위 항목의 prefix 로.
|
||||
if let Some((mut first_key, _)) = doc.as_table_mut().iter_mut().next() {
|
||||
first_key.leaf_decor_mut().set_prefix(format!("{HEADER}\n"));
|
||||
@@ -413,6 +424,112 @@ pub fn step_3_to_4(doc: &mut DocumentMut, changes: &mut Vec<MigrationChange>) {
|
||||
});
|
||||
}
|
||||
|
||||
/// v5: `[ingest.image.ocr]` 의 12개 **엔진** 키를 새 공유 블록 `[ingest.ocr]`
|
||||
/// 로 끌어올린다(`enabled` 은 미디어별 토글이라 제외 — 끌어올리면 공유 블록이
|
||||
/// pdf 의 `enabled` 까지 켜버려 동작이 바뀐다). image 는 unique 필드가 없으므로
|
||||
/// 공유 블록의 canonical source 로 삼는다. `[ingest.pdf.ocr]` 은 손대지 않는다
|
||||
/// — reconcile 이 pdf 의 **concrete-default 키**를 명시 상태로 채우므로(아래
|
||||
/// `resolve_ocr` 의 "미디어가 명시한 키는 공유 overlay 가 덮지 않음" 규칙에 의해)
|
||||
/// 그 키들의 pdf effective 값은 불변. **단 Option 키(endpoint/det_model/rec_model/
|
||||
/// dict, default `None`)는 reconcile 가 pdf 에 채우지 않으므로**(None 은
|
||||
/// annotated-default 에서 누락) image-only 인 채 끌어올리면 overlay 가 pdf 로
|
||||
/// 누출된다 — `OPTION_OCR_KEYS` 가드로 막는다. 멱등: 키가 이미 옮겨졌으면 no-op.
|
||||
///
|
||||
/// 결과적으로 image 의 effective OCR 엔진 설정은 `[ingest.ocr]` 에서, pdf 는
|
||||
/// 자신의 (reconcile 로 채워진 concrete 키 + image-local 로 남은 Option 키) 기준으로
|
||||
/// resolve 되어 양쪽 모두 마이그레이션 전 값을 유지 — `ingest_config_signature` 불변.
|
||||
const SHARED_OCR_ENGINE_KEYS: [&str; 12] = [
|
||||
"engine",
|
||||
"model",
|
||||
"endpoint",
|
||||
"languages",
|
||||
"max_pixels",
|
||||
"request_timeout_secs",
|
||||
"det_model",
|
||||
"rec_model",
|
||||
"dict",
|
||||
"score_thresh",
|
||||
"unclip_ratio",
|
||||
"max_boxes",
|
||||
];
|
||||
|
||||
/// `SHARED_OCR_ENGINE_KEYS` 중 default 가 `None` 인 Option-typed 키. reconcile 가
|
||||
/// pdf 블록에 채우지 않으므로(None 은 annotated-default 에서 누락) image-only 인 채
|
||||
/// 공유 블록으로 끌어올리면 `resolve_ocr` overlay 가 pdf 로 누출시킨다 — pdf 가
|
||||
/// 명시한 경우에만 hoist(그때는 pdf 의 명시값이 overlay 를 막는다).
|
||||
const OPTION_OCR_KEYS: [&str; 4] = ["endpoint", "det_model", "rec_model", "dict"];
|
||||
|
||||
pub fn step_4_to_5(doc: &mut DocumentMut, changes: &mut Vec<MigrationChange>) {
|
||||
// image OCR 블록이 없으면 끌어올릴 게 없음(reconcile 이 빈 `[ingest.ocr]`
|
||||
// 를 추가). 멱등 진입점.
|
||||
let img_present = doc
|
||||
.get("ingest")
|
||||
.and_then(|i| i.get("image"))
|
||||
.and_then(|i| i.get("ocr"))
|
||||
.and_then(Item::as_table)
|
||||
.is_some();
|
||||
if !img_present {
|
||||
return;
|
||||
}
|
||||
|
||||
let mut lifted_any = false;
|
||||
for key in SHARED_OCR_ENGINE_KEYS {
|
||||
// image.ocr 에 key 가 있고, ingest.ocr 에 아직 없으면 통째(decor 포함) 이동.
|
||||
let has_in_image = doc
|
||||
.get("ingest")
|
||||
.and_then(|i| i.get("image"))
|
||||
.and_then(|i| i.get("ocr"))
|
||||
.and_then(Item::as_table)
|
||||
.is_some_and(|t| t.contains_key(key));
|
||||
if !has_in_image {
|
||||
continue;
|
||||
}
|
||||
// Option 키는 pdf 가 명시한 경우에만 끌어올린다 — image-only Option 키를
|
||||
// 공유 블록에 올리면 resolve_ocr overlay 가 (medium_has(pdf,key)=false 라)
|
||||
// pdf 로 누출돼 pdf 의 effective OCR(endpoint/asset 경로)가 v4 와 달라지고,
|
||||
// det/rec/dict 는 ingest_config_signature 에 들어가 강제 재색인까지 유발한다.
|
||||
// 자세한 이유는 OPTION_OCR_KEYS 주석 참조.
|
||||
if OPTION_OCR_KEYS.contains(&key) {
|
||||
let pdf_has_key = doc
|
||||
.get("ingest")
|
||||
.and_then(|i| i.get("pdf"))
|
||||
.and_then(|i| i.get("ocr"))
|
||||
.and_then(Item::as_table)
|
||||
.is_some_and(|t| t.contains_key(key));
|
||||
if !pdf_has_key {
|
||||
continue;
|
||||
}
|
||||
}
|
||||
let already_in_shared = doc
|
||||
.get("ingest")
|
||||
.and_then(|i| i.get("ocr"))
|
||||
.and_then(Item::as_table)
|
||||
.is_some_and(|t| t.contains_key(key));
|
||||
if already_in_shared {
|
||||
// 공유 블록에 이미 있으면 image 쪽 중복 키는 그냥 제거(공유가 우선).
|
||||
if let Some(img) = doc["ingest"]["image"]["ocr"].as_table_mut() {
|
||||
img.remove(key);
|
||||
}
|
||||
continue;
|
||||
}
|
||||
move_table(
|
||||
doc,
|
||||
&["ingest", "image", "ocr", key],
|
||||
&["ingest", "ocr", key],
|
||||
changes,
|
||||
);
|
||||
lifted_any = true;
|
||||
}
|
||||
|
||||
if lifted_any {
|
||||
changes.push(MigrationChange {
|
||||
kind: ChangeKind::AddedSection,
|
||||
path: "ingest.ocr".to_string(),
|
||||
detail: "OCR 엔진 키를 공유 [ingest.ocr] 로 통합(image/pdf 중복 제거)".to_string(),
|
||||
});
|
||||
}
|
||||
}
|
||||
|
||||
/// 파일의 schema_version(없으면 1) 부터 CURRENT 까지 step 적용.
|
||||
fn run_steps(doc: &mut DocumentMut, from: u32, changes: &mut Vec<MigrationChange>) {
|
||||
if from < 2 {
|
||||
@@ -424,6 +541,9 @@ fn run_steps(doc: &mut DocumentMut, from: u32, changes: &mut Vec<MigrationChange
|
||||
if from < 4 {
|
||||
step_3_to_4(doc, changes);
|
||||
}
|
||||
if from < 5 {
|
||||
step_4_to_5(doc, changes);
|
||||
}
|
||||
}
|
||||
|
||||
/// 사용자 config.toml 텍스트를 받아 step 체인 + reconciliation + version
|
||||
@@ -477,6 +597,21 @@ pub fn migrate_document(text: &str) -> MigrationOutcome {
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
/// v5: parse a config text and run the shared-OCR resolution the way
|
||||
/// `Config::from_file` does (overlay `[ingest.ocr]` down into the
|
||||
/// per-medium concrete blocks), then clear the now-applied shared block.
|
||||
/// The result is the canonical *effective* config — comparable against
|
||||
/// `Config::defaults()` regardless of whether the engine knobs live in
|
||||
/// the shared block (annotated default doc) or the per-medium blocks
|
||||
/// (in-memory `defaults()`).
|
||||
fn parse_effective(text: &str) -> crate::Config {
|
||||
let parsed = toml::from_str::<toml::Value>(text).ok();
|
||||
let mut cfg: crate::Config = toml::from_str(text).expect("parse config");
|
||||
cfg.resolve_ocr(parsed.as_ref());
|
||||
cfg.ingest.ocr = crate::SharedOcrEngineCfg::default();
|
||||
cfg
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn annotated_default_has_per_key_comments() {
|
||||
let text = annotated_default_document().to_string();
|
||||
@@ -487,9 +622,9 @@ mod tests {
|
||||
text.contains("paddle-onnx 는 번들 모델"),
|
||||
"ocr.model 주석 누락:\n{text}"
|
||||
);
|
||||
// 주석 추가가 파싱을 깨지 않는다.
|
||||
let back: crate::Config = toml::from_str(&text).expect("parse annotated default");
|
||||
assert_eq!(back, crate::Config::defaults());
|
||||
// 주석 추가가 파싱을 깨지 않고, v5 OCR resolution 후 effective 값이
|
||||
// defaults 와 동일.
|
||||
assert_eq!(parse_effective(&text), crate::Config::defaults());
|
||||
}
|
||||
|
||||
#[test]
|
||||
@@ -498,10 +633,11 @@ mod tests {
|
||||
let text = doc.to_string();
|
||||
// v3: 미디어 형식 섹션이 전부 `[ingest.*]` 하위로 통합됐다. IngestCfg
|
||||
// 는 스칼라(병렬도) 필드가 있어 bare `[ingest]` + 하위 테이블이 함께
|
||||
// 직렬화된다.
|
||||
// 직렬화된다. v5: 공유 `[ingest.ocr]` 엔진 블록이 추가됐다.
|
||||
for section in [
|
||||
"[workspace]",
|
||||
"[ingest]",
|
||||
"[ingest.ocr]",
|
||||
"[ingest.chunking]",
|
||||
"[ingest.code]",
|
||||
"[ingest.image.ocr]",
|
||||
@@ -512,8 +648,8 @@ mod tests {
|
||||
assert!(text.contains(section), "missing {section}:\n{text}");
|
||||
}
|
||||
assert!(text.contains("# "), "no comments attached");
|
||||
let back: crate::Config = toml::from_str(&text).expect("parse annotated default");
|
||||
assert_eq!(back, crate::Config::defaults());
|
||||
// v5: effective 값(공유 OCR resolution 후)이 defaults 와 동일.
|
||||
assert_eq!(parse_effective(&text), crate::Config::defaults());
|
||||
}
|
||||
|
||||
#[test]
|
||||
@@ -739,7 +875,7 @@ root = \"/my/notes\"
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn migrate_document_v3_to_v4_adds_sources_and_is_idempotent() {
|
||||
fn migrate_document_v3_to_current_adds_sources_and_is_idempotent() {
|
||||
let v3 = "\
|
||||
schema_version = 3
|
||||
|
||||
@@ -749,10 +885,10 @@ exclude = []
|
||||
";
|
||||
let outcome = migrate_document(v3);
|
||||
assert_eq!(outcome.from_schema_version, 3);
|
||||
assert_eq!(outcome.to_schema_version, 4);
|
||||
assert_eq!(outcome.to_schema_version, CURRENT_SCHEMA_VERSION);
|
||||
assert!(outcome.changed());
|
||||
assert!(outcome.new_text.contains("[[workspace.sources]]"));
|
||||
assert_eq!(read_schema_version(&outcome.new_text), 4);
|
||||
assert_eq!(read_schema_version(&outcome.new_text), CURRENT_SCHEMA_VERSION);
|
||||
let again = migrate_document(&outcome.new_text);
|
||||
assert!(!again.changed(), "not idempotent: {:?}", again.changes);
|
||||
assert_eq!(again.new_text, outcome.new_text);
|
||||
@@ -765,4 +901,170 @@ exclude = []
|
||||
assert_eq!(outcome.from_schema_version, 1);
|
||||
assert_eq!(read_schema_version(&outcome.new_text), CURRENT_SCHEMA_VERSION);
|
||||
}
|
||||
|
||||
/// v4 → v5 무손실 라운드트립: `[ingest.image.ocr]` / `[ingest.pdf.ocr]` 이
|
||||
/// 채워진 v4 config 을 마이그레이션 → from_file 로 로드(공유 OCR resolution
|
||||
/// 포함) → image/pdf OCR 의 effective 값이 마이그레이션 전과 정확히 동일해야
|
||||
/// 한다. image 의 비-default engine(paddle-onnx) 이 공유 블록으로 끌어올려진
|
||||
/// 뒤에도 보존되는지(이전 버그) + pdf 의 고유 값(qwen 모델·2048px)이 공유
|
||||
/// overlay 에 오염되지 않는지를 함께 검증한다.
|
||||
#[test]
|
||||
fn migrate_v4_to_v5_preserves_effective_ocr() {
|
||||
let v4 = "\
|
||||
schema_version = 4
|
||||
|
||||
[workspace]
|
||||
root = \"/my/notes\"
|
||||
exclude = []
|
||||
|
||||
[[workspace.sources]]
|
||||
id = \"default\"
|
||||
root = \"/my/notes\"
|
||||
|
||||
[ingest.image.ocr]
|
||||
enabled = true
|
||||
engine = \"paddle-onnx\"
|
||||
model = \"gemma4:e4b\"
|
||||
languages = [\"eng\", \"kor\"]
|
||||
max_pixels = 1280
|
||||
request_timeout_secs = 450
|
||||
det_model = \"/custom/det.onnx\"
|
||||
score_thresh = 0.45
|
||||
|
||||
[ingest.pdf.ocr]
|
||||
enabled = true
|
||||
always_on = false
|
||||
engine = \"ollama-vision\"
|
||||
model = \"qwen2.5vl:7b\"
|
||||
languages = [\"eng\", \"kor\"]
|
||||
max_pixels = 2048
|
||||
request_timeout_secs = 240
|
||||
valid_ratio_threshold = 0.6
|
||||
min_char_count = 25
|
||||
lang_hint = \"kor\"
|
||||
";
|
||||
// pre-migration effective values, loaded the v4 way (no shared block).
|
||||
let dir = std::env::temp_dir().join(format!("kebab_v5_rt_{}", std::process::id()));
|
||||
std::fs::create_dir_all(&dir).unwrap();
|
||||
let p4 = dir.join("v4.toml");
|
||||
std::fs::write(&p4, v4).unwrap();
|
||||
let before = crate::Config::from_file(&p4).expect("load v4");
|
||||
|
||||
// migrate → load the v5 text via from_file (runs resolve_ocr).
|
||||
let outcome = migrate_document(v4);
|
||||
assert_eq!(outcome.from_schema_version, 4);
|
||||
assert_eq!(outcome.to_schema_version, 5);
|
||||
assert!(outcome.changed());
|
||||
assert!(
|
||||
outcome.new_text.contains("[ingest.ocr]"),
|
||||
"shared block missing:\n{}",
|
||||
outcome.new_text
|
||||
);
|
||||
let p5 = dir.join("v5.toml");
|
||||
std::fs::write(&p5, &outcome.new_text).unwrap();
|
||||
let after = crate::Config::from_file(&p5).expect("load v5");
|
||||
|
||||
// image OCR effective 값 전부 보존(특히 비-default engine paddle-onnx).
|
||||
assert_eq!(after.image_ocr().enabled, before.image_ocr().enabled);
|
||||
assert_eq!(after.image_ocr().engine, "paddle-onnx");
|
||||
assert_eq!(after.image_ocr().engine, before.image_ocr().engine);
|
||||
assert_eq!(after.image_ocr().model, before.image_ocr().model);
|
||||
assert_eq!(after.image_ocr().languages, before.image_ocr().languages);
|
||||
assert_eq!(after.image_ocr().max_pixels, 1280);
|
||||
assert_eq!(after.image_ocr().max_pixels, before.image_ocr().max_pixels);
|
||||
assert_eq!(
|
||||
after.image_ocr().request_timeout_secs,
|
||||
before.image_ocr().request_timeout_secs
|
||||
);
|
||||
assert_eq!(
|
||||
after.image_ocr().det_model.as_deref(),
|
||||
Some("/custom/det.onnx")
|
||||
);
|
||||
assert_eq!(after.image_ocr().det_model, before.image_ocr().det_model);
|
||||
assert!((after.image_ocr().score_thresh - 0.45).abs() < 1e-6);
|
||||
assert_eq!(after.image_ocr(), before.image_ocr());
|
||||
|
||||
// pdf OCR effective 값 전부 보존(공유 overlay 가 image 값으로 오염 X).
|
||||
assert_eq!(after.pdf_ocr().engine, "ollama-vision");
|
||||
assert_eq!(after.pdf_ocr().model, "qwen2.5vl:7b");
|
||||
assert_eq!(after.pdf_ocr().max_pixels, 2048);
|
||||
// image-only Option 키(det_model)는 pdf 로 누출되지 않아야 한다. v4 바이너리
|
||||
// 시맨틱(pdf 가 det_model 미선언 → None)을 **명시적으로** 검증 — `after ==
|
||||
// before` 만으론 둘 다 같은 (잠재 오염) 파이프라인을 거쳐 tautology 가 된다.
|
||||
assert_eq!(
|
||||
after.pdf_ocr().det_model,
|
||||
None,
|
||||
"image-only det_model leaked into pdf (v4 binary resolves None)"
|
||||
);
|
||||
assert_eq!(after.pdf_ocr(), before.pdf_ocr());
|
||||
|
||||
// 멱등.
|
||||
let again = migrate_document(&outcome.new_text);
|
||||
assert!(!again.changed(), "v5 재실행 변경: {:?}", again.changes);
|
||||
assert_eq!(again.new_text, outcome.new_text);
|
||||
}
|
||||
|
||||
/// 회귀(PR #214 리뷰): image 가 Option-typed OCR 키(endpoint/det_model)를
|
||||
/// 설정하고 pdf 는 미설정인 비대칭 v4 config. 마이그레이션 후 그 키가 공유
|
||||
/// 블록을 거쳐 pdf 로 누출되면 pdf OCR 가 image endpoint/asset 으로 잘못
|
||||
/// 라우팅되고 det/rec/dict 는 강제 재색인을 유발한다. v4 바이너리 시맨틱은
|
||||
/// pdf=None 이어야 하고, image 는 자기 값을 유지해야 한다.
|
||||
#[test]
|
||||
fn migrate_v4_to_v5_image_only_option_keys_stay_image_local() {
|
||||
let v4 = "\
|
||||
schema_version = 4
|
||||
|
||||
[workspace]
|
||||
root = \"/n\"
|
||||
exclude = []
|
||||
|
||||
[[workspace.sources]]
|
||||
id = \"default\"
|
||||
root = \"/n\"
|
||||
|
||||
[ingest.image.ocr]
|
||||
enabled = true
|
||||
engine = \"ollama-vision\"
|
||||
endpoint = \"http://image-ocr-host:9999\"
|
||||
det_model = \"/custom/det.onnx\"
|
||||
|
||||
[ingest.pdf.ocr]
|
||||
enabled = true
|
||||
engine = \"ollama-vision\"
|
||||
model = \"qwen2.5vl:7b\"
|
||||
";
|
||||
let outcome = migrate_document(v4);
|
||||
assert_eq!(outcome.to_schema_version, 5);
|
||||
|
||||
let dir = std::env::temp_dir().join(format!("kebab_v5_opt_{}", std::process::id()));
|
||||
std::fs::create_dir_all(&dir).unwrap();
|
||||
let p5 = dir.join("v5.toml");
|
||||
std::fs::write(&p5, &outcome.new_text).unwrap();
|
||||
let after = crate::Config::from_file(&p5).expect("load v5");
|
||||
|
||||
// image 는 자기 Option 값을 유지.
|
||||
assert_eq!(
|
||||
after.image_ocr().endpoint.as_deref(),
|
||||
Some("http://image-ocr-host:9999")
|
||||
);
|
||||
assert_eq!(
|
||||
after.image_ocr().det_model.as_deref(),
|
||||
Some("/custom/det.onnx")
|
||||
);
|
||||
// pdf 는 누출 없이 None (v4 바이너리 시맨틱).
|
||||
assert_eq!(
|
||||
after.pdf_ocr().endpoint,
|
||||
None,
|
||||
"image-only endpoint leaked into pdf"
|
||||
);
|
||||
assert_eq!(
|
||||
after.pdf_ocr().det_model,
|
||||
None,
|
||||
"image-only det_model leaked into pdf"
|
||||
);
|
||||
|
||||
// 멱등.
|
||||
let again = migrate_document(&outcome.new_text);
|
||||
assert!(!again.changed(), "v5 재실행 변경: {:?}", again.changes);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -81,7 +81,6 @@ default_k = 10
|
||||
hybrid_fusion = "rrf"
|
||||
rrf_k = 60
|
||||
snippet_chars = 220
|
||||
cache_capacity = 256
|
||||
stale_threshold_days = 30
|
||||
|
||||
[rag]
|
||||
|
||||
@@ -11,9 +11,13 @@ const USER_V2: &str = include_str!("fixtures/user_v2_config.toml");
|
||||
fn user_v2_migrates_losslessly() {
|
||||
let out = migrate_document(USER_V2);
|
||||
assert_eq!(out.from_schema_version, 2);
|
||||
// v2 → CURRENT(=4): v3 의 [ingest.*] relocation 에 더해 v4 의
|
||||
// [[workspace.sources]] default source 미러링까지 적용된다.
|
||||
assert_eq!(out.to_schema_version, 4);
|
||||
// v2 → CURRENT(=5): v3 의 [ingest.*] relocation, v4 의
|
||||
// [[workspace.sources]] default source 미러링, v5 의 공유 [ingest.ocr]
|
||||
// 통합까지 적용된다.
|
||||
assert_eq!(
|
||||
out.to_schema_version,
|
||||
kebab_config::migrate::CURRENT_SCHEMA_VERSION
|
||||
);
|
||||
let t = &out.new_text;
|
||||
|
||||
// 사용자 값 보존.
|
||||
@@ -36,15 +40,23 @@ fn user_v2_migrates_losslessly() {
|
||||
assert!(!t.contains("\n[image.ocr]"));
|
||||
assert!(!t.contains("\n[indexing]"));
|
||||
|
||||
// v3 Config 로 parse + 값 동일.
|
||||
let cfg: kebab_config::Config = toml::from_str(t).expect("v3 parse");
|
||||
assert!(cfg.ingest.image.ocr.enabled);
|
||||
assert_eq!(cfg.ingest.image.ocr.engine, "paddle-onnx");
|
||||
// v5: 공유 [ingest.ocr] 통합 후 image 엔진 키는 공유 블록에 산다.
|
||||
assert!(t.contains("[ingest.ocr]"), "공유 OCR 블록 누락:\n{t}");
|
||||
|
||||
// effective 값은 from_file(공유 OCR resolution 포함)로 검증한다 —
|
||||
// image 의 engine=paddle-onnx 가 공유 블록으로 끌어올려진 뒤에도 보존.
|
||||
let dir = std::env::temp_dir().join(format!("kebab_mv3_{}", std::process::id()));
|
||||
std::fs::create_dir_all(&dir).unwrap();
|
||||
let p = dir.join("config.toml");
|
||||
std::fs::write(&p, t).unwrap();
|
||||
let cfg = kebab_config::Config::from_file(&p).expect("v5 from_file");
|
||||
assert!(cfg.image_ocr().enabled);
|
||||
assert_eq!(cfg.image_ocr().engine, "paddle-onnx");
|
||||
assert_eq!(cfg.models.embedding.model, "snowflake-arctic-embed2");
|
||||
assert_eq!(cfg.models.llm.endpoint, "http://192.168.0.2:11943");
|
||||
// pdf paddle 값 보존(v2 비대칭 → pdf 대칭 키로 복사). user 의 pdf.ocr 는
|
||||
// engine=paddle-onnx 이고 자체 det_model 없으므로 번들(None) 유지.
|
||||
assert_eq!(cfg.ingest.pdf.ocr.engine, "paddle-onnx");
|
||||
assert_eq!(cfg.pdf_ocr().engine, "paddle-onnx");
|
||||
|
||||
// 멱등.
|
||||
let again = migrate_document(t);
|
||||
|
||||
@@ -63,15 +63,17 @@ fn pdf_ocr_defaults_off_with_qwen_3b() {
|
||||
assert_eq!(cfg.ingest.pdf.ocr.lang_hint.as_deref(), Some("kor"));
|
||||
}
|
||||
|
||||
// Test 3: env var override — 4 keys 의 typical override case.
|
||||
// Test 3: env var override — pdf enabled + shared engine knob.
|
||||
// v5: `model` moved to the shared `KEBAB_OCR_MODEL` (sets both mediums);
|
||||
// `enabled` stays per-medium addressable.
|
||||
// KEBAB_PDF_OCR_ALWAYS_ON and KEBAB_PDF_OCR_VALID_RATIO_THRESHOLD removed
|
||||
// from env surface (config-only) in Unit 2 surface trim.
|
||||
#[test]
|
||||
fn pdf_ocr_env_overrides() {
|
||||
let mut env: HashMap<String, String> = HashMap::new();
|
||||
env.insert("KEBAB_PDF_OCR_ENABLED".to_string(), "true".to_string());
|
||||
env.insert(
|
||||
"KEBAB_PDF_OCR_MODEL".to_string(),
|
||||
"qwen2.5vl:7b".to_string(),
|
||||
);
|
||||
env.insert("KEBAB_OCR_MODEL".to_string(), "qwen2.5vl:7b".to_string());
|
||||
// always_on and valid_ratio_threshold arms gone — silently ignored below.
|
||||
env.insert("KEBAB_PDF_OCR_ALWAYS_ON".to_string(), "true".to_string());
|
||||
env.insert(
|
||||
"KEBAB_PDF_OCR_VALID_RATIO_THRESHOLD".to_string(),
|
||||
@@ -82,8 +84,10 @@ fn pdf_ocr_env_overrides() {
|
||||
|
||||
assert!(cfg.ingest.pdf.ocr.enabled);
|
||||
assert_eq!(cfg.ingest.pdf.ocr.model, "qwen2.5vl:7b");
|
||||
assert!(cfg.ingest.pdf.ocr.always_on);
|
||||
assert!((cfg.ingest.pdf.ocr.valid_ratio_threshold - 0.75).abs() < 1e-6);
|
||||
// always_on arm gone — stays at default (false).
|
||||
assert!(!cfg.ingest.pdf.ocr.always_on);
|
||||
// valid_ratio_threshold arm gone — stays at default (0.5).
|
||||
assert!((cfg.ingest.pdf.ocr.valid_ratio_threshold - 0.5).abs() < 1e-6);
|
||||
|
||||
// 다른 env var 가 default 보존
|
||||
assert_eq!(cfg.ingest.pdf.ocr.engine, "ollama-vision");
|
||||
|
||||
@@ -20,15 +20,6 @@ pub struct Answer {
|
||||
pub usage: TokenUsage,
|
||||
#[serde(with = "time::serde::rfc3339")]
|
||||
pub created_at: OffsetDateTime,
|
||||
/// p9-fb-15: same conversation 의 turn 들이 공유. CLI single-shot
|
||||
/// (history 없음) / TUI 첫 turn 은 None. blake3 해시 또는 사용자
|
||||
/// 명시 (`kebab ask --session <id>`, p9-fb-18).
|
||||
#[serde(default, skip_serializing_if = "Option::is_none")]
|
||||
pub conversation_id: Option<String>,
|
||||
/// p9-fb-15: 같은 conversation 안 0-based 순서. 첫 turn = 0. None
|
||||
/// 이면 single-shot.
|
||||
#[serde(default, skip_serializing_if = "Option::is_none")]
|
||||
pub turn_index: Option<u32>,
|
||||
/// p9-fb-41: multi-hop hop trace. `None` for single-pass asks.
|
||||
/// Each entry records one hop (`decompose` / `decide` / `synthesize`)
|
||||
/// — the LLM call category, the sub-queries emitted, retrieval
|
||||
@@ -73,19 +64,6 @@ pub struct AnswerCitation {
|
||||
pub stale: bool,
|
||||
}
|
||||
|
||||
/// p9-fb-15: history 가 prompt 에 들어갈 때의 한 turn. RAG facade 가
|
||||
/// `Vec<Turn>` 받아 system + history + retrieval + new question 으로
|
||||
/// prompt 빌드. token budget 안에 fit 안 되면 oldest turn 부터 drop
|
||||
/// (newest 우선 보존).
|
||||
#[derive(Clone, Debug, PartialEq, Serialize, Deserialize)]
|
||||
pub struct Turn {
|
||||
pub question: String,
|
||||
pub answer: String,
|
||||
pub citations: Vec<AnswerCitation>,
|
||||
#[serde(with = "time::serde::rfc3339")]
|
||||
pub created_at: OffsetDateTime,
|
||||
}
|
||||
|
||||
/// p9-fb-41: one entry in [`Answer::hops`] — the per-iteration trace
|
||||
/// of a multi-hop ask. The pipeline appends a `HopRecord` per LLM
|
||||
/// call (decompose / decide / synthesize) so a `--multi-hop` user
|
||||
@@ -275,8 +253,6 @@ mod tests {
|
||||
latency_ms: 0,
|
||||
},
|
||||
created_at: datetime!(2026-05-09 12:00:00 UTC),
|
||||
conversation_id: None,
|
||||
turn_index: None,
|
||||
hops: None,
|
||||
verification: None,
|
||||
};
|
||||
|
||||
@@ -31,7 +31,7 @@ pub mod versions;
|
||||
|
||||
pub use answer::{
|
||||
Answer, AnswerCitation, AnswerRetrievalSummary, HopKind, HopRecord, ModelRef, RefusalReason,
|
||||
TokenUsage, TraceId, Turn, VerificationSummary,
|
||||
TokenUsage, TraceId, VerificationSummary,
|
||||
};
|
||||
pub use asset::{AssetStorage, RawAsset, SourceUri, WorkspacePath};
|
||||
pub use chunk::Chunk;
|
||||
@@ -60,7 +60,7 @@ pub use search::{
|
||||
SearchQuery, SearchTrace, TraceCandidate, TraceFusionInput, TraceTiming,
|
||||
};
|
||||
pub use traits::{
|
||||
ChatSessionRepo, ChatSessionRow, ChatTurnRow, ChunkPolicy, Chunker, DocumentStore, Embedder,
|
||||
ChunkPolicy, Chunker, DocumentStore, Embedder,
|
||||
EmbeddingInput, EmbeddingKind, ExtractConfig, ExtractContext, Extractor, FinishReason,
|
||||
GenerateRequest, JobRepo, LanguageModel, Retriever, SourceConnector, SourceScope, TokenChunk,
|
||||
VectorStore,
|
||||
|
||||
@@ -39,6 +39,19 @@ pub struct ExtractContext<'a> {
|
||||
pub asset: &'a RawAsset,
|
||||
pub workspace_root: &'a Path,
|
||||
pub config: &'a ExtractConfig,
|
||||
/// `[[workspace.sources]]`: id of the source this asset is being
|
||||
/// ingested from. The markdown extractor threads it into `BodyHints`
|
||||
/// so `parse_frontmatter` stamps `Metadata.source_id` (frontmatter
|
||||
/// does not override it). Other extractors (pdf / image / code) leave
|
||||
/// it `None` and the kebab-app handler stamps `source_id` post-extract.
|
||||
pub source_id: Option<&'a str>,
|
||||
/// `[[workspace.sources]]`: per-source default `trust_level`. The
|
||||
/// markdown extractor threads it into `BodyHints` so `parse_frontmatter`
|
||||
/// can apply the precedence chain (frontmatter > this default >
|
||||
/// hardcoded `Primary`) *inside* extraction — which is why it must be
|
||||
/// carried here rather than stamped after. `None` for non-markdown
|
||||
/// extractors (their frontmatter never carries a `trust_level`).
|
||||
pub source_trust: Option<crate::metadata::TrustLevel>,
|
||||
}
|
||||
|
||||
#[derive(Clone, Debug, PartialEq, Serialize, Deserialize)]
|
||||
@@ -232,67 +245,3 @@ pub trait JobRepo {
|
||||
fn list(&self, filter: &JobFilter) -> anyhow::Result<Vec<JobRow>>;
|
||||
}
|
||||
|
||||
// ── p9-fb-17: chat session persistence ────────────────────────────────
|
||||
|
||||
/// Persistent multi-turn chat session — header row in `chat_sessions`.
|
||||
/// Per-turn rows live in `chat_turns` (see [`ChatTurnRow`]).
|
||||
#[derive(Clone, Debug, PartialEq, Serialize, Deserialize)]
|
||||
pub struct ChatSessionRow {
|
||||
pub session_id: String,
|
||||
/// Unix epoch seconds at session creation time.
|
||||
pub created_at: i64,
|
||||
/// Unix epoch seconds, bumped on every `append_turn`.
|
||||
pub updated_at: i64,
|
||||
/// Optional human-readable label — defaults to the first
|
||||
/// question's first ~40 chars on creation.
|
||||
pub title: Option<String>,
|
||||
/// Snapshot of `prompt_template_version`, `llm.model`,
|
||||
/// `max_context_tokens`, etc. — same shape as
|
||||
/// `eval_runs.config_snapshot_json`. JSON string so the schema
|
||||
/// can grow without an SQLite ALTER.
|
||||
pub config_snapshot_json: String,
|
||||
}
|
||||
|
||||
/// One Q/A pair inside a `ChatSessionRow`. `turn_index` is monotonic
|
||||
/// per session (0-based).
|
||||
#[derive(Clone, Debug, PartialEq, Serialize, Deserialize)]
|
||||
pub struct ChatTurnRow {
|
||||
/// `blake3(session_id || turn_index)` (32 hex). Stable per (session,
|
||||
/// turn) so a re-append at the same index is rejected via PK.
|
||||
pub turn_id: String,
|
||||
pub session_id: String,
|
||||
pub turn_index: u32,
|
||||
pub question: String,
|
||||
pub answer: String,
|
||||
/// `Vec<Citation>` JSON-encoded so a session resume can replay
|
||||
/// the same citation markers the user saw originally.
|
||||
pub citations_json: String,
|
||||
pub created_at: i64,
|
||||
}
|
||||
|
||||
/// Persistence trait for multi-turn chat sessions. Implemented by
|
||||
/// `kebab-store-sqlite::SqliteStore`; consumed by `kebab-app` and the
|
||||
/// future CLI / TUI session UIs (p9-fb-18).
|
||||
pub trait ChatSessionRepo {
|
||||
/// Create a new session. `session_id` is caller-supplied — auto
|
||||
/// derivation lives in `kebab-app`. Errors on PK collision.
|
||||
fn create_session(&self, row: &ChatSessionRow) -> anyhow::Result<()>;
|
||||
|
||||
/// Look up a session by id; `Ok(None)` when missing.
|
||||
fn get_session(&self, session_id: &str) -> anyhow::Result<Option<ChatSessionRow>>;
|
||||
|
||||
/// Most-recent-updated-first list of sessions, capped at `limit`.
|
||||
fn list_sessions(&self, limit: usize) -> anyhow::Result<Vec<ChatSessionRow>>;
|
||||
|
||||
/// Delete a session and (CASCADE) every turn under it.
|
||||
fn delete_session(&self, session_id: &str) -> anyhow::Result<()>;
|
||||
|
||||
/// Append a turn at `turn.turn_index`. Bumps the parent's
|
||||
/// `updated_at`. PK collision (same session_id + turn_index) is
|
||||
/// an error — the caller assigns the next monotonic index.
|
||||
fn append_turn(&self, turn: &ChatTurnRow) -> anyhow::Result<()>;
|
||||
|
||||
/// All turns for `session_id`, ordered by `turn_index ASC`.
|
||||
/// Empty vec when the session has no turns yet.
|
||||
fn list_turns(&self, session_id: &str) -> anyhow::Result<Vec<ChatTurnRow>>;
|
||||
}
|
||||
|
||||
@@ -1,50 +0,0 @@
|
||||
[package]
|
||||
name = "kebab-embed-candle"
|
||||
version = { workspace = true }
|
||||
edition = { workspace = true }
|
||||
rust-version = { workspace = true }
|
||||
license = { workspace = true }
|
||||
repository = { workspace = true }
|
||||
description = "Pure-Rust candle adapter implementing kb_core::Embedder (multilingual-e5-large, NUMA-safe thread cap)"
|
||||
|
||||
[dependencies]
|
||||
kebab-core = { path = "../kebab-core" }
|
||||
kebab-config = { path = "../kebab-config" }
|
||||
# candle stack — pinned to the workspace-locked crates.io release (0.10.x),
|
||||
# same versions the Phase 0 spike compiled so build artifacts are reused.
|
||||
candle-core = "0.10.2"
|
||||
candle-nn = "0.10.2"
|
||||
candle-transformers = "0.10.2"
|
||||
tokenizers = "0.21"
|
||||
hf-hub = { version = "0.4", features = ["ureq"] }
|
||||
serde_json = { workspace = true }
|
||||
# Thread cap: a one-shot global rayon pool sizes candle's CPU threads
|
||||
# (the Phase 0 spike proved RAYON_NUM_THREADS caps candle), so a NUMA host
|
||||
# can keep onnxruntime's hard-coded 48-intra-op heap corruption at bay.
|
||||
rayon = "1"
|
||||
anyhow = { workspace = true }
|
||||
tracing = { workspace = true }
|
||||
|
||||
[features]
|
||||
# opt-in: run candle on the Apple Silicon GPU (Metal). macOS-only — the build
|
||||
# enables candle's metal backend and `select_device()` picks Metal (CPU fallback
|
||||
# on failure). Lets an M-series Mac ingest e5-large on GPU (10×+ vs CPU); the
|
||||
# resulting vectors are cross-compatible with the CPU path (same model), so the
|
||||
# Linux server can serve queries on CPU candle.
|
||||
metal = ["candle-core/metal", "candle-nn/metal", "candle-transformers/metal"]
|
||||
|
||||
[dev-dependencies]
|
||||
# Integration-test binaries can only see the library's public API + these,
|
||||
# not the library's own (non-dev) dependencies — so rayon/kebab-config/kebab-core
|
||||
# are repeated here for tests/parity.rs and tests/thread_cap.rs.
|
||||
kebab-embed-local = { path = "../kebab-embed-local" }
|
||||
# arctic↔Ollama parity test drives the real Ollama adapter for the reference
|
||||
# vectors (tests/arctic_ollama_parity.rs, `#[ignore]` — live Ollama).
|
||||
kebab-embed-ollama = { path = "../kebab-embed-ollama" }
|
||||
kebab-config = { path = "../kebab-config" }
|
||||
kebab-core = { path = "../kebab-core" }
|
||||
rayon = "1"
|
||||
tempfile = { workspace = true }
|
||||
|
||||
[lints]
|
||||
workspace = true
|
||||
@@ -1,619 +0,0 @@
|
||||
//! `kebab-embed-candle` — [`CandleEmbedder`], a pure-Rust (candle)
|
||||
//! implementation of [`Embedder`](kebab_core::Embedder).
|
||||
//!
|
||||
//! Runs an XLM-RoBERTa-large embedding model through `candle`
|
||||
//! (`candle-transformers`' XLM-RoBERTa) instead of onnxruntime. Two models
|
||||
//! are wired through a small **registry** ([`MODEL_REGISTRY`]):
|
||||
//!
|
||||
//! * `multilingual-e5-large` — the same weights the default
|
||||
//! [`FastembedEmbedder`](kebab_embed_local) uses (mean pooling,
|
||||
//! `query: `/`passage: ` prefixes). candle is the NUMA-safe drop-in:
|
||||
//! fastembed 4.9's onnxruntime hard-codes 48 intra-op threads, which
|
||||
//! corrupts the heap (double-free) on dual-socket NUMA hosts. candle's
|
||||
//! CPU backend sizes its threads off the global rayon pool, so a one-shot
|
||||
//! [`rayon::ThreadPoolBuilder`] cap (config `num_threads` / env
|
||||
//! `KEBAB_EMBED_THREADS`) keeps the worker count NUMA-safe.
|
||||
//! * `snowflake-arctic-embed-l-v2.0` — Snowflake's arctic-embed v2.0
|
||||
//! (CLS pooling, `query: ` on queries / no prefix on documents). Same
|
||||
//! XLM-RoBERTa-large architecture, dim 1024, so it rides the exact same
|
||||
//! tokenize → forward → L2 pipeline; only the pooling step and prefixes
|
||||
//! differ (both keyed off the per-model [`EmbedModelSpec`]).
|
||||
//!
|
||||
//! Output parity with the onnxruntime path (for e5) was proven by the
|
||||
//! Phase 0 spike (cosine 1.000000); the arctic path's pooling/prefix
|
||||
//! correctness is pinned by an `#[ignore]`d cosine>0.99 cross-check against
|
||||
//! Ollama's `snowflake-arctic-embed2` (see `tests/arctic_ollama_parity.rs`).
|
||||
//! The shared pipeline:
|
||||
//!
|
||||
//! 1. instruction prefix per [`EmbedModelSpec`] (query/doc);
|
||||
//! 2. tokenize (max_len 512, batch-longest padding, special tokens);
|
||||
//! 3. XLM-RoBERTa forward on the selected [`Device`];
|
||||
//! 4. pooling — mean (attention-mask-weighted) or CLS (first token);
|
||||
//! 5. L2 normalization.
|
||||
//!
|
||||
//! Model files (`config.json`, `tokenizer.json`, `model.safetensors`) are
|
||||
//! fetched via `hf-hub` into `{config.storage.model_dir}/candle/` (hf-hub's
|
||||
//! cache layout namespaces by repo, so e5 and arctic never collide).
|
||||
//!
|
||||
//! This crate is **opt-in** (`config.models.embedding.provider = "candle"`);
|
||||
//! the default provider stays `fastembed`. See
|
||||
//! `docs/superpowers/specs/2026-06-01-embed-candle-track-spec.md` and
|
||||
//! `docs/superpowers/specs/2026-06-03-arctic-embedder-spec.md`.
|
||||
|
||||
use std::sync::Mutex;
|
||||
|
||||
use anyhow::{Context, Result};
|
||||
use candle_core::{DType, Device, Tensor};
|
||||
use candle_nn::VarBuilder;
|
||||
use candle_transformers::models::xlm_roberta::{Config as XlmConfig, XLMRobertaModel};
|
||||
use kebab_config::{Config, expand_path};
|
||||
use kebab_core::{Embedder, EmbeddingInput, EmbeddingKind, EmbeddingModelId, EmbeddingVersion};
|
||||
use tokenizers::{PaddingParams, PaddingStrategy, Tokenizer, TruncationParams};
|
||||
|
||||
/// Subdirectory under `config.storage.model_dir` where the candle adapter
|
||||
/// caches safetensors + tokenizer. Mirrors `kebab-embed-local`'s
|
||||
/// `fastembed/` subdir so the two backends never collide.
|
||||
const CANDLE_CACHE_SUBDIR: &str = "candle";
|
||||
|
||||
/// Token truncation length (both e5 and arctic-embed-l-v2.0 train at 512).
|
||||
const MAX_LEN: usize = 512;
|
||||
|
||||
/// Env var that overrides `config.models.embedding.num_threads`. Read once in
|
||||
/// [`CandleEmbedder::new`]; `0`/unset/unparseable means "leave rayon default".
|
||||
const ENV_EMBED_THREADS: &str = "KEBAB_EMBED_THREADS";
|
||||
|
||||
/// Pooling strategy over the model's last hidden state. Keyed per-model by
|
||||
/// [`EmbedModelSpec::pooling`] — e5 is mean, arctic is CLS.
|
||||
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
|
||||
pub enum Pooling {
|
||||
/// Attention-mask-weighted mean over all tokens (e5 / sentence-transformers
|
||||
/// `pooling_mode_mean_tokens`).
|
||||
Mean,
|
||||
/// First token (`<s>`/`[CLS]`) hidden state (arctic-embed v2.0 —
|
||||
/// `1_Pooling/config.json` has `pooling_mode_cls_token: true`).
|
||||
Cls,
|
||||
}
|
||||
|
||||
/// One supported embedding model: the HF repo candle downloads, the pooling
|
||||
/// strategy, and the e5-style instruction prefixes. [`MODEL_REGISTRY`] maps a
|
||||
/// `config.models.embedding.model` value to one of these.
|
||||
#[derive(Clone, Copy, Debug)]
|
||||
pub struct EmbedModelSpec {
|
||||
/// The short `config.models.embedding.model` value that selects this spec.
|
||||
pub name: &'static str,
|
||||
/// HuggingFace repo id candle fetches `config.json` / `tokenizer.json` /
|
||||
/// `model.safetensors` from.
|
||||
pub hf_repo: &'static str,
|
||||
/// Pooling over the last hidden state.
|
||||
pub pooling: Pooling,
|
||||
/// Prefix prepended to **query** inputs before tokenization.
|
||||
pub query_prefix: &'static str,
|
||||
/// Prefix prepended to **document** inputs before tokenization (arctic
|
||||
/// uses `""` — documents are embedded raw).
|
||||
pub doc_prefix: &'static str,
|
||||
/// Expected embedding dimension (model hidden size).
|
||||
pub dim: usize,
|
||||
/// Suffix folded into `model_version` so switching **to** this model
|
||||
/// triggers the `embedding_version` cascade even if the operator forgets
|
||||
/// to bump `config.version`. `None` keeps the bare `config.version` — used
|
||||
/// by e5 so candle-e5 and fastembed-e5 report the *same* version and stay
|
||||
/// interchangeable (the NUMA drop-in invariant — Phase 0 cosine 1.0).
|
||||
pub version_tag: Option<&'static str>,
|
||||
}
|
||||
|
||||
/// The models the candle adapter can load. Adding a model = one entry here
|
||||
/// (plus, for a non-XLM-R architecture, a new forward path — both current
|
||||
/// entries are XLM-RoBERTa-large so they share everything but pooling/prefix).
|
||||
static MODEL_REGISTRY: &[EmbedModelSpec] = &[
|
||||
EmbedModelSpec {
|
||||
name: "multilingual-e5-large",
|
||||
hf_repo: "intfloat/multilingual-e5-large",
|
||||
pooling: Pooling::Mean,
|
||||
query_prefix: "query: ",
|
||||
doc_prefix: "passage: ",
|
||||
dim: 1024,
|
||||
version_tag: None,
|
||||
},
|
||||
EmbedModelSpec {
|
||||
name: "snowflake-arctic-embed-l-v2.0",
|
||||
hf_repo: "Snowflake/snowflake-arctic-embed-l-v2.0",
|
||||
pooling: Pooling::Cls,
|
||||
query_prefix: "query: ",
|
||||
doc_prefix: "",
|
||||
dim: 1024,
|
||||
version_tag: Some("arctic-cls"),
|
||||
},
|
||||
];
|
||||
|
||||
/// Look up a model spec by `config.models.embedding.model`. Accepts either the
|
||||
/// short `name` or the full `hf_repo` id (mirrors the old e5 guard, which
|
||||
/// accepted both `multilingual-e5-large` and `intfloat/multilingual-e5-large`).
|
||||
pub(crate) fn lookup_spec(model: &str) -> Option<&'static EmbedModelSpec> {
|
||||
MODEL_REGISTRY
|
||||
.iter()
|
||||
.find(|s| s.name == model || s.hf_repo == model)
|
||||
}
|
||||
|
||||
/// Comma-separated list of supported model names, for the
|
||||
/// unsupported-model error message.
|
||||
fn supported_models() -> String {
|
||||
MODEL_REGISTRY
|
||||
.iter()
|
||||
.map(|s| s.name)
|
||||
.collect::<Vec<_>>()
|
||||
.join("`, `")
|
||||
}
|
||||
|
||||
/// Pure-Rust candle adapter. Construct via [`CandleEmbedder::new`]; the
|
||||
/// constructor downloads the model on first use, so share one instance.
|
||||
pub struct CandleEmbedder {
|
||||
// candle's `forward` is `&self`, but `XLMRobertaModel` is not guaranteed
|
||||
// `Sync`; the `Mutex` both supplies that bound and serializes inference
|
||||
// (callers batch sequentially anyway — same rationale as
|
||||
// `FastembedEmbedder`).
|
||||
model: Mutex<XLMRobertaModel>,
|
||||
tokenizer: Tokenizer,
|
||||
device: Device,
|
||||
/// The resolved model spec (pooling + prefixes) — drives `embed` and
|
||||
/// `embed_batch`.
|
||||
spec: &'static EmbedModelSpec,
|
||||
model_id: EmbeddingModelId,
|
||||
version: EmbeddingVersion,
|
||||
dimensions: usize,
|
||||
batch_size: usize,
|
||||
}
|
||||
|
||||
impl CandleEmbedder {
|
||||
/// Build an embedder from `Config`. Resolves the model spec from
|
||||
/// `config.models.embedding.model`, applies the NUMA thread cap, fetches
|
||||
/// the model into `{model_dir}/candle/`, and validates that the model's
|
||||
/// hidden size matches `config.models.embedding.dimensions` before
|
||||
/// returning.
|
||||
pub fn new(config: &Config) -> Result<Self> {
|
||||
// 1. NUMA thread cap. env `KEBAB_EMBED_THREADS` wins over the config
|
||||
// field; `0`/unset leaves rayon's default. `build_global` errors if
|
||||
// the pool was already initialized — intentionally ignored so a
|
||||
// second embedder (or a prior rayon user) is a no-op, not a failure.
|
||||
let n_threads = std::env::var(ENV_EMBED_THREADS)
|
||||
.ok()
|
||||
.and_then(|v| v.parse::<usize>().ok())
|
||||
.unwrap_or(config.models.embedding.num_threads as usize);
|
||||
if n_threads > 0 {
|
||||
if apply_thread_cap(n_threads) {
|
||||
tracing::info!(
|
||||
target: "kebab-embed-candle",
|
||||
num_threads = n_threads,
|
||||
"capped global rayon pool for candle CPU backend"
|
||||
);
|
||||
} else {
|
||||
tracing::debug!(
|
||||
target: "kebab-embed-candle",
|
||||
requested = n_threads,
|
||||
"global rayon pool already initialized; thread cap not applied"
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
// 1b. Model registry lookup. If the operator configured a model the
|
||||
// candle adapter doesn't know, fail fast (BEFORE the ~2GB
|
||||
// download) — never silently download one model and then label its
|
||||
// vectors with another name via `model_id()`, which would mislabel
|
||||
// `embedding_version` and corrupt a mixed index.
|
||||
let want = config.models.embedding.model.as_str();
|
||||
let spec = lookup_spec(want).ok_or_else(|| {
|
||||
anyhow::anyhow!(
|
||||
"candle provider supports the models `{}`, but \
|
||||
config.models.embedding.model = '{want}'. Use provider=fastembed \
|
||||
for other models, or pick a supported one.",
|
||||
supported_models()
|
||||
)
|
||||
})?;
|
||||
|
||||
// 2. Resolve `{data_dir}/models/candle/` exactly like the fastembed
|
||||
// adapter resolves its own subdir.
|
||||
let data_dir = expand_path(&config.storage.data_dir, "");
|
||||
let model_dir = expand_path(&config.storage.model_dir, &data_dir.to_string_lossy());
|
||||
let cache_dir = model_dir.join(CANDLE_CACHE_SUBDIR);
|
||||
std::fs::create_dir_all(&cache_dir)
|
||||
.with_context(|| format!("create candle cache dir {}", cache_dir.display()))?;
|
||||
|
||||
let device = select_device();
|
||||
|
||||
// 3. Fetch model files via hf-hub into the candle cache.
|
||||
tracing::info!(
|
||||
target: "kebab-embed-candle",
|
||||
cache_dir = %cache_dir.display(),
|
||||
model = spec.hf_repo,
|
||||
pooling = ?spec.pooling,
|
||||
"loading candle embedding model (first run downloads ~2GB safetensors)"
|
||||
);
|
||||
let api = hf_hub::api::sync::ApiBuilder::new()
|
||||
.with_cache_dir(cache_dir.clone())
|
||||
.build()
|
||||
.context("kb-embed-candle: build hf-hub api")?;
|
||||
let repo = api.model(spec.hf_repo.to_string());
|
||||
let config_path = repo.get("config.json").context("download config.json")?;
|
||||
let tokenizer_path = repo
|
||||
.get("tokenizer.json")
|
||||
.context("download tokenizer.json")?;
|
||||
let weights_path = repo
|
||||
.get("model.safetensors")
|
||||
.context("download model.safetensors")?;
|
||||
|
||||
// 4. Build the candle XLM-RoBERTa model.
|
||||
let cfg_json = std::fs::read_to_string(&config_path)
|
||||
.with_context(|| format!("read {}", config_path.display()))?;
|
||||
let cfg: XlmConfig =
|
||||
serde_json::from_str(&cfg_json).context("kb-embed-candle: parse XLM-R config")?;
|
||||
|
||||
// Validate dim BEFORE building the model so a misconfigured
|
||||
// `dimensions` fails cheaply (matches FastembedEmbedder's contract).
|
||||
check_dim(cfg.hidden_size, config.models.embedding.dimensions)?;
|
||||
|
||||
let vb = unsafe {
|
||||
VarBuilder::from_mmaped_safetensors(&[weights_path], DType::F32, &device)
|
||||
.context("kb-embed-candle: mmap safetensors")?
|
||||
};
|
||||
let model =
|
||||
XLMRobertaModel::new(&cfg, vb).context("kb-embed-candle: build XLMRobertaModel")?;
|
||||
|
||||
let mut tokenizer = Tokenizer::from_file(&tokenizer_path)
|
||||
.map_err(|e| anyhow::anyhow!("kb-embed-candle: load tokenizer: {e}"))?;
|
||||
tokenizer
|
||||
.with_padding(Some(PaddingParams {
|
||||
strategy: PaddingStrategy::BatchLongest,
|
||||
..Default::default()
|
||||
}))
|
||||
.with_truncation(Some(TruncationParams {
|
||||
max_length: MAX_LEN,
|
||||
..Default::default()
|
||||
}))
|
||||
.map_err(|e| anyhow::anyhow!("kb-embed-candle: set truncation: {e}"))?;
|
||||
|
||||
// model_version: fold the model tag in for non-e5 models so a switch
|
||||
// triggers the embedding_version cascade; e5 keeps the bare
|
||||
// config.version to stay interchangeable with fastembed-e5.
|
||||
let version = match spec.version_tag {
|
||||
Some(tag) => {
|
||||
EmbeddingVersion(format!("{}+{}", config.models.embedding.version, tag))
|
||||
}
|
||||
None => EmbeddingVersion(config.models.embedding.version.clone()),
|
||||
};
|
||||
|
||||
tracing::info!(
|
||||
target: "kebab-embed-candle",
|
||||
dimensions = cfg.hidden_size,
|
||||
layers = cfg.num_hidden_layers,
|
||||
model = spec.name,
|
||||
"candle embedding model loaded"
|
||||
);
|
||||
|
||||
Ok(Self {
|
||||
model: Mutex::new(model),
|
||||
tokenizer,
|
||||
device,
|
||||
spec,
|
||||
model_id: EmbeddingModelId(config.models.embedding.model.clone()),
|
||||
version,
|
||||
dimensions: cfg.hidden_size,
|
||||
batch_size: config.models.embedding.batch_size.max(1),
|
||||
})
|
||||
}
|
||||
|
||||
/// Embed one batch of **already-prefixed** strings (the per-model prefix
|
||||
/// is applied by the caller [`CandleEmbedder::embed`]) through the candle
|
||||
/// pipeline: tokenize → forward → pool (mean|CLS) → L2.
|
||||
fn embed_batch(&self, prefixed: &[String]) -> Result<Vec<Vec<f32>>> {
|
||||
let encodings = self
|
||||
.tokenizer
|
||||
.encode_batch(prefixed.to_vec(), true)
|
||||
.map_err(|e| anyhow::anyhow!("kb-embed-candle: encode_batch: {e}"))?;
|
||||
|
||||
let bsz = encodings.len();
|
||||
// `embed` already returns early on empty input and `.chunks()` never
|
||||
// yields an empty slice, so this is currently unreachable — but guard
|
||||
// the index so a future refactor can't turn it into a panic.
|
||||
let Some(first) = encodings.first() else {
|
||||
return Ok(Vec::new());
|
||||
};
|
||||
let seq = first.get_ids().len();
|
||||
|
||||
let mut ids = Vec::with_capacity(bsz * seq);
|
||||
let mut mask = Vec::with_capacity(bsz * seq);
|
||||
for enc in &encodings {
|
||||
ids.extend(enc.get_ids().iter().copied());
|
||||
mask.extend(enc.get_attention_mask().iter().map(|&m| m as f32));
|
||||
}
|
||||
|
||||
let input_ids = Tensor::from_vec(ids, (bsz, seq), &self.device)?;
|
||||
let attn_f32 = Tensor::from_vec(mask, (bsz, seq), &self.device)?;
|
||||
let token_type_ids = input_ids.zeros_like()?;
|
||||
|
||||
let hidden = {
|
||||
let guard = self
|
||||
.model
|
||||
.lock()
|
||||
.unwrap_or_else(std::sync::PoisonError::into_inner);
|
||||
// forward: (input_ids, attention_mask, token_type_ids, past,
|
||||
// encoder_hidden, encoder_mask)
|
||||
guard.forward(&input_ids, &attn_f32, &token_type_ids, None, None, None)?
|
||||
};
|
||||
|
||||
// Pooling — per the model spec.
|
||||
let pooled = match self.spec.pooling {
|
||||
Pooling::Mean => {
|
||||
// attention-mask-weighted mean pooling
|
||||
let mask3 = attn_f32.unsqueeze(2)?; // (b, seq, 1)
|
||||
let summed = hidden.broadcast_mul(&mask3)?.sum(1)?; // (b, hidden)
|
||||
// counts ≥ 1 always: every input is prefixed AND special
|
||||
// tokens are added (encode_batch(_, true)), so no row has an
|
||||
// all-zero mask. If that invariant ever breaks, broadcast_div
|
||||
// would emit NaN vectors.
|
||||
let counts = mask3.sum(1)?; // (b, 1)
|
||||
summed.broadcast_div(&counts)?
|
||||
}
|
||||
Pooling::Cls => {
|
||||
// CLS pooling: the first token's hidden state. arctic-embed
|
||||
// v2.0 prepends `<s>` (the XLM-R BOS/CLS) at index 0, so
|
||||
// `hidden[:, 0, :]` is the sentence embedding.
|
||||
hidden.narrow(1, 0, 1)?.squeeze(1)? // (b, hidden)
|
||||
}
|
||||
};
|
||||
|
||||
// L2 normalize
|
||||
let norm = pooled.sqr()?.sum_keepdim(1)?.sqrt()?;
|
||||
let normalized = pooled.broadcast_div(&norm)?;
|
||||
|
||||
// `.contiguous()` before host copy: broadcast ops can leave a strided
|
||||
// view, which `to_vec2` rejects on the Metal backend (CPU tolerates it).
|
||||
Ok(normalized.contiguous()?.to_vec2::<f32>()?)
|
||||
}
|
||||
}
|
||||
|
||||
impl Embedder for CandleEmbedder {
|
||||
fn model_id(&self) -> EmbeddingModelId {
|
||||
self.model_id.clone()
|
||||
}
|
||||
|
||||
fn model_version(&self) -> EmbeddingVersion {
|
||||
self.version.clone()
|
||||
}
|
||||
|
||||
fn dimensions(&self) -> usize {
|
||||
self.dimensions
|
||||
}
|
||||
|
||||
fn embed(&self, inputs: &[EmbeddingInput<'_>]) -> Result<Vec<Vec<f32>>> {
|
||||
if inputs.is_empty() {
|
||||
return Ok(Vec::new());
|
||||
}
|
||||
|
||||
// Per-model instruction prefix BEFORE tokenization (same convention as
|
||||
// FastembedEmbedder for e5; arctic uses `query: `/no-prefix).
|
||||
let prefixed: Vec<String> = inputs.iter().map(|i| prefix_input(self.spec, i)).collect();
|
||||
|
||||
let mut out: Vec<Vec<f32>> = Vec::with_capacity(prefixed.len());
|
||||
for chunk in prefixed.chunks(self.batch_size) {
|
||||
let batch = self.embed_batch(chunk)?;
|
||||
for v in &batch {
|
||||
if v.len() != self.dimensions {
|
||||
anyhow::bail!(
|
||||
"candle returned vector of length {} but adapter expects {}",
|
||||
v.len(),
|
||||
self.dimensions
|
||||
);
|
||||
}
|
||||
}
|
||||
out.extend(batch);
|
||||
}
|
||||
|
||||
debug_assert_eq!(out.len(), inputs.len());
|
||||
Ok(out)
|
||||
}
|
||||
}
|
||||
|
||||
/// Build the prefixed string for one [`EmbeddingInput`] using the model spec.
|
||||
/// Free function so a unit test can pin the format without loading the model.
|
||||
/// For e5 this is byte-identical to `kebab-embed-local`'s `prefix_input` — the
|
||||
/// two backends MUST agree there or their vectors diverge.
|
||||
fn prefix_input(spec: &EmbedModelSpec, input: &EmbeddingInput<'_>) -> String {
|
||||
match input.kind {
|
||||
EmbeddingKind::Document => format!("{}{}", spec.doc_prefix, input.text),
|
||||
EmbeddingKind::Query => format!("{}{}", spec.query_prefix, input.text),
|
||||
}
|
||||
}
|
||||
|
||||
/// Select the compute device. Built with the `metal` feature (Apple Silicon
|
||||
/// GPU), try Metal and fall back to CPU on failure; otherwise CPU. Metal only
|
||||
/// compiles/runs on macOS — the Linux server builds the CPU path. Embedding
|
||||
/// vectors are model-defined, so Metal-produced and CPU-produced embeddings
|
||||
/// are cross-compatible (a Mac can ingest on GPU, the server query on CPU).
|
||||
fn select_device() -> Device {
|
||||
#[cfg(feature = "metal")]
|
||||
{
|
||||
match Device::new_metal(0) {
|
||||
Ok(d) => {
|
||||
tracing::info!(target: "kebab-embed-candle", "candle device = Metal (GPU)");
|
||||
return d;
|
||||
}
|
||||
Err(e) => {
|
||||
tracing::warn!(
|
||||
target: "kebab-embed-candle",
|
||||
error = %e,
|
||||
"Metal device unavailable; falling back to CPU"
|
||||
);
|
||||
}
|
||||
}
|
||||
}
|
||||
tracing::info!(target: "kebab-embed-candle", "candle device = CPU");
|
||||
Device::Cpu
|
||||
}
|
||||
|
||||
/// Apply a one-shot global rayon thread cap (the NUMA-safety lever). Returns
|
||||
/// `true` if this call set the pool, `false` if it was already initialized
|
||||
/// (cap not applied) or `n_threads == 0`. `#[doc(hidden)] pub` so the
|
||||
/// thread-cap test can drive it without loading the 2GB model.
|
||||
#[doc(hidden)]
|
||||
pub fn apply_thread_cap(n_threads: usize) -> bool {
|
||||
if n_threads == 0 {
|
||||
return false;
|
||||
}
|
||||
rayon::ThreadPoolBuilder::new()
|
||||
.num_threads(n_threads)
|
||||
.build_global()
|
||||
.is_ok()
|
||||
}
|
||||
|
||||
/// Compare model hidden size against the configured dim. Extracted so a unit
|
||||
/// test can exercise the error branch without loading the model.
|
||||
pub(crate) fn check_dim(model_dim: usize, cfg_dim: usize) -> Result<()> {
|
||||
if model_dim != cfg_dim {
|
||||
anyhow::bail!(
|
||||
"dimension mismatch: model={model_dim}, config={cfg_dim}; \
|
||||
update `config.models.embedding.dimensions` to match the model \
|
||||
(or pick a different model)."
|
||||
);
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
fn e5_spec() -> &'static EmbedModelSpec {
|
||||
lookup_spec("multilingual-e5-large").expect("e5 in registry")
|
||||
}
|
||||
|
||||
fn arctic_spec() -> &'static EmbedModelSpec {
|
||||
lookup_spec("snowflake-arctic-embed-l-v2.0").expect("arctic in registry")
|
||||
}
|
||||
|
||||
// ── registry ─────────────────────────────────────────────────────
|
||||
|
||||
#[test]
|
||||
fn registry_resolves_e5_by_name_and_hf_repo() {
|
||||
assert_eq!(
|
||||
lookup_spec("multilingual-e5-large").map(|s| s.name),
|
||||
Some("multilingual-e5-large")
|
||||
);
|
||||
assert_eq!(
|
||||
lookup_spec("intfloat/multilingual-e5-large").map(|s| s.name),
|
||||
Some("multilingual-e5-large")
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn registry_resolves_arctic_and_its_pooling_is_cls() {
|
||||
let s = arctic_spec();
|
||||
assert_eq!(s.name, "snowflake-arctic-embed-l-v2.0");
|
||||
assert_eq!(s.hf_repo, "Snowflake/snowflake-arctic-embed-l-v2.0");
|
||||
assert_eq!(s.pooling, Pooling::Cls);
|
||||
assert_eq!(s.dim, 1024);
|
||||
assert_eq!(s.version_tag, Some("arctic-cls"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn registry_e5_is_mean_pooling_no_version_tag() {
|
||||
let s = e5_spec();
|
||||
assert_eq!(s.pooling, Pooling::Mean);
|
||||
assert_eq!(s.version_tag, None);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn registry_rejects_unknown_model() {
|
||||
assert!(lookup_spec("multilingual-e5-small").is_none());
|
||||
}
|
||||
|
||||
// ── prefix_input ─────────────────────────────────────────────────
|
||||
// e5 prefixes MUST match kebab-embed-local::prefix_input or candle vs
|
||||
// fastembed parity breaks; arctic uses query-only prefixing.
|
||||
|
||||
#[test]
|
||||
fn e5_prefix_document_uses_passage() {
|
||||
let input = EmbeddingInput {
|
||||
text: "hello world",
|
||||
kind: EmbeddingKind::Document,
|
||||
};
|
||||
assert_eq!(prefix_input(e5_spec(), &input), "passage: hello world");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn e5_prefix_query_uses_query() {
|
||||
let input = EmbeddingInput {
|
||||
text: "hello world",
|
||||
kind: EmbeddingKind::Query,
|
||||
};
|
||||
assert_eq!(prefix_input(e5_spec(), &input), "query: hello world");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn arctic_prefix_query_uses_query_doc_is_bare() {
|
||||
let doc = EmbeddingInput {
|
||||
text: "후입선출 자료구조",
|
||||
kind: EmbeddingKind::Document,
|
||||
};
|
||||
let qry = EmbeddingInput {
|
||||
text: "스택 자료구조",
|
||||
kind: EmbeddingKind::Query,
|
||||
};
|
||||
// arctic: documents are embedded raw, queries get `query: `.
|
||||
assert_eq!(prefix_input(arctic_spec(), &doc), "후입선출 자료구조");
|
||||
assert_eq!(prefix_input(arctic_spec(), &qry), "query: 스택 자료구조");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn prefix_handles_empty_text() {
|
||||
let doc = EmbeddingInput {
|
||||
text: "",
|
||||
kind: EmbeddingKind::Document,
|
||||
};
|
||||
let qry = EmbeddingInput {
|
||||
text: "",
|
||||
kind: EmbeddingKind::Query,
|
||||
};
|
||||
assert_eq!(prefix_input(e5_spec(), &doc), "passage: ");
|
||||
assert_eq!(prefix_input(e5_spec(), &qry), "query: ");
|
||||
assert_eq!(prefix_input(arctic_spec(), &doc), "");
|
||||
assert_eq!(prefix_input(arctic_spec(), &qry), "query: ");
|
||||
}
|
||||
|
||||
// ── check_dim ────────────────────────────────────────────────────
|
||||
|
||||
#[test]
|
||||
fn check_dim_passes_for_1024() {
|
||||
check_dim(1024, 1024).expect("matching dims must pass");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn check_dim_rejects_384_vs_1024() {
|
||||
let err = check_dim(384, 1024).expect_err("dim mismatch must error");
|
||||
let msg = format!("{err}");
|
||||
assert!(
|
||||
msg.contains("384") && msg.contains("1024"),
|
||||
"error must mention both dims, got: {msg}"
|
||||
);
|
||||
}
|
||||
|
||||
// ── model guard ──────────────────────────────────────────────────
|
||||
// A model name not in the registry must fail fast (BEFORE the ~2GB
|
||||
// download), so we never download one model yet label its vectors with
|
||||
// another name via model_id() — which would mislabel embedding_version.
|
||||
|
||||
#[test]
|
||||
fn new_rejects_unsupported_model() {
|
||||
let mut config = kebab_config::Config::defaults();
|
||||
config.models.embedding.model = "multilingual-e5-small".to_string();
|
||||
// num_threads defaults to 0, so no global rayon side effect here.
|
||||
// `.err()` (not `expect_err`) avoids requiring `CandleEmbedder: Debug`
|
||||
// — it holds a Mutex/Tokenizer and intentionally derives no Debug.
|
||||
let err = CandleEmbedder::new(&config)
|
||||
.err()
|
||||
.expect("unsupported model must error");
|
||||
let msg = format!("{err:#}");
|
||||
assert!(
|
||||
msg.contains("candle provider supports the models"),
|
||||
"expected model-registry error, got: {msg}"
|
||||
);
|
||||
}
|
||||
}
|
||||
@@ -1,128 +0,0 @@
|
||||
//! arctic-embed-l-v2.0 correctness gate (`#[ignore]` — needs the ~2GB candle
|
||||
//! model + a live Ollama serving `snowflake-arctic-embed2`).
|
||||
//!
|
||||
//! This is the load-bearing pooling/prefix check for the arctic integration.
|
||||
//! The recall measurement that justified adopting arctic (recall@10 130/132)
|
||||
//! went through Ollama's `snowflake-arctic-embed2`. The candle path
|
||||
//! re-implements the model (XLM-RoBERTa-large + **CLS** pooling + `query: ` on
|
||||
//! queries / no prefix on documents). If candle's pooling or prefix is wrong,
|
||||
//! its vectors silently diverge from the measured route and the 130 number
|
||||
//! does NOT carry over. This test pins them together: per-sentence cosine
|
||||
//! between the candle vector and the Ollama vector must be **> 0.99**.
|
||||
//!
|
||||
//! `#[ignore]` because it depends on an external Ollama daemon (CI is
|
||||
//! headless/offline). The leader MUST run it once before merge.
|
||||
//!
|
||||
//! ## Manual run
|
||||
//!
|
||||
//! 1. Confirm Ollama is reachable and has the model:
|
||||
//! ```sh
|
||||
//! curl -s http://192.168.0.47:11434/api/tags # should list snowflake-arctic-embed2
|
||||
//! ```
|
||||
//! 2. Run (downloads the ~2GB candle safetensors on first run):
|
||||
//! ```sh
|
||||
//! CARGO_TARGET_DIR=/build/out/cargo-target \
|
||||
//! KEBAB_ARCTIC_OLLAMA_ENDPOINT=http://192.168.0.47:11434 \
|
||||
//! cargo test -p kebab-embed-candle --test arctic_ollama_parity -- --ignored --nocapture
|
||||
//! ```
|
||||
//! The endpoint defaults to `http://192.168.0.47:11434` if the env is unset.
|
||||
//!
|
||||
//! Record the printed `ARCTIC_PARITY_SUMMARY cosine_min=...` in
|
||||
//! `/tmp/arctic-result.md` + `tasks/HOTFIXES.md`.
|
||||
|
||||
use kebab_config::Config;
|
||||
use kebab_core::{Embedder, EmbeddingInput, EmbeddingKind};
|
||||
use kebab_embed_candle::CandleEmbedder;
|
||||
use kebab_embed_ollama::OllamaEmbedder;
|
||||
|
||||
const DOGFOOD_CONFIG: &str = "/build/dogfood/config.toml";
|
||||
const DEFAULT_OLLAMA_ENDPOINT: &str = "http://192.168.0.47:11434";
|
||||
|
||||
/// Mixed Korean / English + the descriptive-recall shapes arctic was adopted
|
||||
/// for (synonym / abbreviation / English term). Covers both prefix paths.
|
||||
const SENTENCES: &[&str] = &[
|
||||
"스택 자료구조",
|
||||
"후입선출 방식으로 동작하는 자료구조",
|
||||
"큐는 선입선출 자료구조이다",
|
||||
"Rust ownership and the borrow checker",
|
||||
"소유권과 빌림 검사기는 메모리 안전성을 보장한다",
|
||||
"SVM 은 support vector machine 의 약자이다",
|
||||
"정렬 알고리즘의 시간 복잡도",
|
||||
"The capital of France is Paris.",
|
||||
];
|
||||
|
||||
fn cosine(a: &[f32], b: &[f32]) -> f32 {
|
||||
let dot: f32 = a.iter().zip(b).map(|(x, y)| x * y).sum();
|
||||
let na: f32 = a.iter().map(|x| x * x).sum::<f32>().sqrt();
|
||||
let nb: f32 = b.iter().map(|x| x * x).sum::<f32>().sqrt();
|
||||
dot / (na * nb)
|
||||
}
|
||||
|
||||
/// Base config: prefer the canonical dogfood config (for storage/cache roots),
|
||||
/// fall back to `Config::defaults()` so the test still runs on a bare clone.
|
||||
fn base_config() -> Config {
|
||||
Config::load(Some(std::path::Path::new(DOGFOOD_CONFIG))).unwrap_or_else(|_| Config::defaults())
|
||||
}
|
||||
|
||||
#[test]
|
||||
#[ignore = "needs ~2GB candle model + live Ollama (snowflake-arctic-embed2); run manually before merge"]
|
||||
fn candle_arctic_matches_ollama_arctic() {
|
||||
let endpoint = std::env::var("KEBAB_ARCTIC_OLLAMA_ENDPOINT")
|
||||
.unwrap_or_else(|_| DEFAULT_OLLAMA_ENDPOINT.to_string());
|
||||
|
||||
// candle side: the in-process arctic model.
|
||||
let mut candle_cfg = base_config();
|
||||
candle_cfg.models.embedding.provider = "candle".to_string();
|
||||
candle_cfg.models.embedding.model = "snowflake-arctic-embed-l-v2.0".to_string();
|
||||
candle_cfg.models.embedding.dimensions = 1024;
|
||||
|
||||
// Ollama side: the reference route the recall numbers came from.
|
||||
let mut ollama_cfg = base_config();
|
||||
ollama_cfg.models.embedding.provider = "ollama".to_string();
|
||||
ollama_cfg.models.embedding.model = "snowflake-arctic-embed2".to_string();
|
||||
ollama_cfg.models.embedding.dimensions = 1024;
|
||||
ollama_cfg.models.embedding.endpoint = Some(endpoint.clone());
|
||||
|
||||
let candle = CandleEmbedder::new(&candle_cfg).expect("build candle arctic embedder");
|
||||
let ollama = OllamaEmbedder::new(&ollama_cfg).expect("build ollama arctic embedder");
|
||||
|
||||
// Exercise BOTH prefix paths so a query-side divergence can't hide.
|
||||
let inputs: Vec<EmbeddingInput> = SENTENCES
|
||||
.iter()
|
||||
.flat_map(|s| {
|
||||
[EmbeddingKind::Document, EmbeddingKind::Query]
|
||||
.into_iter()
|
||||
.map(move |kind| EmbeddingInput { text: s, kind })
|
||||
})
|
||||
.collect();
|
||||
|
||||
let cv = candle.embed(&inputs).expect("candle embed");
|
||||
let ov = ollama
|
||||
.embed(&inputs)
|
||||
.expect("ollama embed (is snowflake-arctic-embed2 pulled @ the endpoint?)");
|
||||
|
||||
assert_eq!(cv.len(), ov.len(), "embedding counts must match");
|
||||
assert_eq!(cv.len(), inputs.len(), "one vector per input");
|
||||
assert_eq!(candle.dimensions(), 1024);
|
||||
|
||||
let mut min_cos = f32::INFINITY;
|
||||
for (i, inp) in inputs.iter().enumerate() {
|
||||
assert_eq!(cv[i].len(), 1024, "candle dim");
|
||||
assert_eq!(ov[i].len(), 1024, "ollama dim");
|
||||
let c = cosine(&cv[i], &ov[i]);
|
||||
min_cos = min_cos.min(c);
|
||||
let kind = match inp.kind {
|
||||
EmbeddingKind::Document => "doc",
|
||||
EmbeddingKind::Query => "qry",
|
||||
};
|
||||
let preview: String = inp.text.chars().take(36).collect();
|
||||
println!("[{i:>2}] {kind} cos={c:.6} {preview}");
|
||||
}
|
||||
|
||||
println!("ARCTIC_PARITY_SUMMARY cosine_min={min_cos:.6} endpoint={endpoint}");
|
||||
assert!(
|
||||
min_cos > 0.99,
|
||||
"candle arctic vs Ollama arctic cosine_min={min_cos:.6} ≤ 0.99 — \
|
||||
pooling/prefix mismatch; the recall=130 measurement will NOT reproduce"
|
||||
);
|
||||
}
|
||||
@@ -1,96 +0,0 @@
|
||||
//! Parity test (spec §7, `#[ignore]` — needs the ~2GB model + network).
|
||||
//!
|
||||
//! Confirms the candle backend reproduces the onnxruntime `FastembedEmbedder`
|
||||
//! vectors closely enough that no re-index is required (spec D-reindex):
|
||||
//! per-sentence cosine ≥ 0.9999, and reports the dimension-wise max absolute
|
||||
//! difference (the number the re-index decision hangs on).
|
||||
//!
|
||||
//! Run manually:
|
||||
//! CARGO_TARGET_DIR=/build/out/cargo-target/target \
|
||||
//! cargo test -p kebab-embed-candle --release -- --ignored --nocapture
|
||||
//!
|
||||
//! Uses the canonical dogfood config so both backends resolve the same model
|
||||
//! identifiers and cache roots.
|
||||
|
||||
use kebab_config::Config;
|
||||
use kebab_core::{Embedder, EmbeddingInput, EmbeddingKind};
|
||||
use kebab_embed_candle::CandleEmbedder;
|
||||
use kebab_embed_local::FastembedEmbedder;
|
||||
|
||||
const DOGFOOD_CONFIG: &str = "/build/dogfood/config.toml";
|
||||
|
||||
/// Mixed Korean / English parity set (≥ 8 sentences, mirrors the Phase 0 spike).
|
||||
const SENTENCES: &[&str] = &[
|
||||
"The quick brown fox jumps over the lazy dog.",
|
||||
"오늘 날씨가 정말 좋아서 산책을 나가고 싶다.",
|
||||
"Rust is a systems programming language focused on safety and performance.",
|
||||
"벡터 검색은 임베딩 사이의 코사인 유사도를 이용한다.",
|
||||
"Machine learning models require large amounts of training data.",
|
||||
"한국어와 영어가 섞인 문장도 멀티링구얼 모델은 잘 처리한다.",
|
||||
"The capital of France is Paris, a city known for its art and culture.",
|
||||
"이 프로젝트는 로컬 우선 지식 베이스와 검색 증강 생성을 목표로 한다.",
|
||||
"Database indexing dramatically speeds up query performance.",
|
||||
"임베딩 모델을 candle 로 옮기면 NUMA 서버에서 안전하게 돌릴 수 있다.",
|
||||
];
|
||||
|
||||
fn cosine(a: &[f32], b: &[f32]) -> f32 {
|
||||
let dot: f32 = a.iter().zip(b).map(|(x, y)| x * y).sum();
|
||||
let na: f32 = a.iter().map(|x| x * x).sum::<f32>().sqrt();
|
||||
let nb: f32 = b.iter().map(|x| x * x).sum::<f32>().sqrt();
|
||||
dot / (na * nb)
|
||||
}
|
||||
|
||||
#[test]
|
||||
#[ignore = "needs ~2GB model + network; run manually for the re-index decision"]
|
||||
fn candle_matches_fastembed() {
|
||||
let config = Config::load(Some(std::path::Path::new(DOGFOOD_CONFIG)))
|
||||
.expect("load dogfood config for parity baseline");
|
||||
|
||||
let candle = CandleEmbedder::new(&config).expect("build CandleEmbedder");
|
||||
let fastembed = FastembedEmbedder::new(&config).expect("build FastembedEmbedder");
|
||||
|
||||
// Cover BOTH prefix paths (`passage:` for Document, `query:` for Query) so
|
||||
// a query-side prefix/pooling divergence can't slip through (reviewer note).
|
||||
let inputs: Vec<EmbeddingInput> = SENTENCES
|
||||
.iter()
|
||||
.flat_map(|s| {
|
||||
[EmbeddingKind::Document, EmbeddingKind::Query]
|
||||
.into_iter()
|
||||
.map(move |kind| EmbeddingInput { text: s, kind })
|
||||
})
|
||||
.collect();
|
||||
|
||||
let cv = candle.embed(&inputs).expect("candle embed");
|
||||
let fv = fastembed.embed(&inputs).expect("fastembed embed");
|
||||
|
||||
assert_eq!(cv.len(), fv.len(), "embedding counts must match");
|
||||
assert_eq!(cv.len(), inputs.len(), "one vector per input");
|
||||
assert_eq!(candle.dimensions(), 1024);
|
||||
|
||||
let mut min_cos = f32::INFINITY;
|
||||
let mut max_abs_diff = 0f32;
|
||||
for (i, inp) in inputs.iter().enumerate() {
|
||||
assert_eq!(cv[i].len(), 1024, "candle dim");
|
||||
assert_eq!(fv[i].len(), 1024, "fastembed dim");
|
||||
let c = cosine(&cv[i], &fv[i]);
|
||||
min_cos = min_cos.min(c);
|
||||
let diff = cv[i]
|
||||
.iter()
|
||||
.zip(&fv[i])
|
||||
.map(|(a, b)| (a - b).abs())
|
||||
.fold(0f32, f32::max);
|
||||
max_abs_diff = max_abs_diff.max(diff);
|
||||
let kind = match inp.kind {
|
||||
EmbeddingKind::Document => "doc",
|
||||
EmbeddingKind::Query => "qry",
|
||||
};
|
||||
let preview: String = inp.text.chars().take(36).collect();
|
||||
println!("[{i:>2}] {kind} cos={c:.6} max_abs_diff={diff:.6e} {preview}");
|
||||
}
|
||||
|
||||
println!("PARITY_SUMMARY cosine_min={min_cos:.6} max_abs_diff={max_abs_diff:.6e}");
|
||||
assert!(
|
||||
min_cos >= 0.9999,
|
||||
"candle vs fastembed cosine_min={min_cos:.6} < 0.9999 — investigate before merge"
|
||||
);
|
||||
}
|
||||
@@ -1,32 +0,0 @@
|
||||
//! Thread-cap test (spec §7). Own integration binary → clean process, so the
|
||||
//! one-shot global rayon pool is initialized exactly once, by us.
|
||||
//!
|
||||
//! Verifies that `apply_thread_cap(4)` sizes the global rayon pool to 4, which
|
||||
//! is the lever that keeps candle's CPU backend NUMA-safe (vs onnxruntime's
|
||||
//! hard-coded 48 intra-op threads).
|
||||
|
||||
use kebab_embed_candle::apply_thread_cap;
|
||||
|
||||
#[test]
|
||||
fn thread_cap_sizes_global_rayon_pool() {
|
||||
// Must run before any other rayon use in this process. As the only test in
|
||||
// this binary that touches rayon, that holds.
|
||||
let applied = apply_thread_cap(4);
|
||||
assert!(applied, "first build_global call should succeed");
|
||||
assert_eq!(
|
||||
rayon::current_num_threads(),
|
||||
4,
|
||||
"global rayon pool must be capped at the requested 4 threads"
|
||||
);
|
||||
|
||||
// A second cap attempt is a no-op (pool already built), not a panic.
|
||||
assert!(
|
||||
!apply_thread_cap(8),
|
||||
"second build_global must report not-applied"
|
||||
);
|
||||
assert_eq!(
|
||||
rayon::current_num_threads(),
|
||||
4,
|
||||
"thread count must stay at the first cap"
|
||||
);
|
||||
}
|
||||
@@ -23,17 +23,18 @@
|
||||
//! See `docs/superpowers/specs/2026-04-27-kebab-final-form-design.md`
|
||||
//! §7.2 (Embedder), §6.4 ([models.embedding]), §9 (versioning).
|
||||
|
||||
use std::path::Path;
|
||||
use std::sync::Mutex;
|
||||
|
||||
use anyhow::{Context, Result};
|
||||
use fastembed::{EmbeddingModel, InitOptions, TextEmbedding};
|
||||
use kebab_config::expand_path;
|
||||
use kebab_config::EmbeddingModelCfg;
|
||||
use kebab_embed::{Embedder, EmbeddingInput, EmbeddingKind, EmbeddingModelId, EmbeddingVersion};
|
||||
|
||||
/// Subdirectory under `config.storage.model_dir` where the fastembed
|
||||
/// adapter writes / reads ONNX + tokenizer files. Hard-coded per task
|
||||
/// spec ("Model files cached under `config.storage.model_dir/fastembed/`").
|
||||
const FASTEMBED_CACHE_SUBDIR: &str = "fastembed";
|
||||
pub const FASTEMBED_CACHE_SUBDIR: &str = "fastembed";
|
||||
|
||||
/// Local fastembed-rs adapter.
|
||||
///
|
||||
@@ -55,37 +56,35 @@ pub struct FastembedEmbedder {
|
||||
}
|
||||
|
||||
impl FastembedEmbedder {
|
||||
/// Build an embedder from `Config`. Validates that
|
||||
/// `config.models.embedding.dimensions` matches the model's actual
|
||||
/// dim BEFORE returning, so a mismatch fails at construction (not on
|
||||
/// first `embed`).
|
||||
pub fn new(config: &kebab_config::Config) -> Result<Self> {
|
||||
// 1. Resolve `{data_dir}/models/fastembed/` from the config
|
||||
// templates. Goes through the shared `kebab_config::expand_path`
|
||||
// so every crate resolves storage paths identically.
|
||||
let data_dir = expand_path(&config.storage.data_dir, "");
|
||||
let model_dir = expand_path(&config.storage.model_dir, &data_dir.to_string_lossy());
|
||||
let cache_dir = model_dir.join(FASTEMBED_CACHE_SUBDIR);
|
||||
std::fs::create_dir_all(&cache_dir)
|
||||
/// Build an embedder from the `[models.embedding]` slice + a resolved
|
||||
/// `cache_dir` (the fastembed subdir under `config.storage.model_dir`;
|
||||
/// the caller resolves it from the storage paths and the
|
||||
/// [`FASTEMBED_CACHE_SUBDIR`] constant). Validates that `cfg.dimensions`
|
||||
/// matches the model's actual dim BEFORE returning, so a mismatch fails
|
||||
/// at construction (not on first `embed`).
|
||||
pub fn new(cfg: &EmbeddingModelCfg, cache_dir: &Path) -> Result<Self> {
|
||||
// 1. The caller resolved `{data_dir}/models/fastembed/`; we own
|
||||
// directory creation so a missing cache dir still works.
|
||||
std::fs::create_dir_all(cache_dir)
|
||||
.with_context(|| format!("create fastembed cache dir {}", cache_dir.display()))?;
|
||||
|
||||
// 2. Resolve the fastembed enum variant from
|
||||
// `config.models.embedding.model`. Currently `multilingual-e5-large`
|
||||
// (default) and `multilingual-e5-small` are wired; other model names
|
||||
// error out with a clear message rather than silently misconfiguring.
|
||||
let model_name = resolve_model(&config.models.embedding.model)?;
|
||||
// 2. Resolve the fastembed enum variant from `cfg.model`. Currently
|
||||
// `multilingual-e5-large` (default) and `multilingual-e5-small`
|
||||
// are wired; other model names error out with a clear message
|
||||
// rather than silently misconfiguring.
|
||||
let model_name = resolve_model(&cfg.model)?;
|
||||
|
||||
// 3. Verify dim match BEFORE loading the model — if the config
|
||||
// is wrong we want to fail without paying the ONNX
|
||||
// initialization cost.
|
||||
let model_info =
|
||||
TextEmbedding::get_model_info(&model_name).context("fastembed: get_model_info")?;
|
||||
check_dim(model_info.dim, config.models.embedding.dimensions)?;
|
||||
check_dim(model_info.dim, cfg.dimensions)?;
|
||||
|
||||
tracing::info!(
|
||||
target: "kebab-embed-local",
|
||||
cache_dir = %cache_dir.display(),
|
||||
model = %config.models.embedding.model,
|
||||
model = %cfg.model,
|
||||
dims = model_info.dim,
|
||||
"initializing FastembedEmbedder"
|
||||
);
|
||||
@@ -95,11 +94,11 @@ impl FastembedEmbedder {
|
||||
// download progress is surfaced via the `tracing::info!`
|
||||
// pair around `TextEmbedding::try_new` instead.
|
||||
let opts = InitOptions::new(model_name.clone())
|
||||
.with_cache_dir(cache_dir.clone())
|
||||
.with_cache_dir(cache_dir.to_path_buf())
|
||||
.with_show_download_progress(false);
|
||||
tracing::info!(
|
||||
target: "kebab-embed-local",
|
||||
model = %config.models.embedding.model,
|
||||
model = %cfg.model,
|
||||
cache_dir = %cache_dir.display(),
|
||||
"loading embedding model (first run downloads model weights — ~470MB for e5-small, ~1.3GB for e5-large)"
|
||||
);
|
||||
@@ -107,17 +106,17 @@ impl FastembedEmbedder {
|
||||
let dimensions = model_info.dim;
|
||||
tracing::info!(
|
||||
target: "kebab-embed-local",
|
||||
model = %config.models.embedding.model,
|
||||
model = %cfg.model,
|
||||
dimensions,
|
||||
"embedding model loaded"
|
||||
);
|
||||
|
||||
Ok(Self {
|
||||
inner: Mutex::new(inner),
|
||||
model_id: EmbeddingModelId(config.models.embedding.model.clone()),
|
||||
version: EmbeddingVersion(config.models.embedding.version.clone()),
|
||||
model_id: EmbeddingModelId(cfg.model.clone()),
|
||||
version: EmbeddingVersion(cfg.version.clone()),
|
||||
dimensions,
|
||||
batch_size: config.models.embedding.batch_size,
|
||||
batch_size: cfg.batch_size,
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
@@ -24,7 +24,15 @@ use std::sync::OnceLock;
|
||||
use std::time::Instant;
|
||||
|
||||
use kebab_embed::{Embedder, EmbeddingInput, EmbeddingKind};
|
||||
use kebab_embed_local::FastembedEmbedder;
|
||||
use kebab_embed_local::{FASTEMBED_CACHE_SUBDIR, FastembedEmbedder};
|
||||
|
||||
/// Resolve the fastembed cache dir from a `Config`'s storage paths,
|
||||
/// mirroring what `kebab-app`'s `embedder()` does at the call site.
|
||||
fn fastembed_cache_dir(cfg: &kebab_config::Config) -> std::path::PathBuf {
|
||||
let data_dir = kebab_config::expand_path(&cfg.storage.data_dir, "");
|
||||
let model_dir = kebab_config::expand_path(&cfg.storage.model_dir, &data_dir.to_string_lossy());
|
||||
model_dir.join(FASTEMBED_CACHE_SUBDIR)
|
||||
}
|
||||
|
||||
/// Build a `Config` whose `data_dir` lives in a per-process temp dir so
|
||||
/// the test never writes into the developer's real `~/.local/share/kebab`.
|
||||
@@ -52,7 +60,8 @@ fn shared_embedder() -> &'static FastembedEmbedder {
|
||||
// and wreck subsequent calls.) The OS will reclaim the leaked
|
||||
// path when the test process exits.
|
||||
let _ = std::mem::ManuallyDrop::new(_tmp);
|
||||
FastembedEmbedder::new(&cfg).expect("init FastembedEmbedder")
|
||||
let cache_dir = fastembed_cache_dir(&cfg);
|
||||
FastembedEmbedder::new(&cfg.models.embedding, &cache_dir).expect("init FastembedEmbedder")
|
||||
})
|
||||
}
|
||||
|
||||
@@ -73,10 +82,11 @@ fn default_config_constructs_with_dims_1024() {
|
||||
fn mismatched_dims_in_config_errors_at_construction() {
|
||||
let (mut cfg, _tmp) = test_config();
|
||||
cfg.models.embedding.dimensions = 512; // model is 1024 (e5-large default)
|
||||
let cache_dir = fastembed_cache_dir(&cfg);
|
||||
// `FastembedEmbedder` deliberately does not implement `Debug`
|
||||
// (its inner ONNX session has no useful debug shape), so we
|
||||
// can't use `expect_err`; match the Result manually.
|
||||
let err = match FastembedEmbedder::new(&cfg) {
|
||||
let err = match FastembedEmbedder::new(&cfg.models.embedding, &cache_dir) {
|
||||
Ok(_) => panic!("dim mismatch must error"),
|
||||
Err(e) => e,
|
||||
};
|
||||
|
||||
@@ -4,11 +4,10 @@
|
||||
//!
|
||||
//! ## Why this exists
|
||||
//!
|
||||
//! The candle backend ([`kebab-embed-candle`]) runs arctic-embed-l-v2.0
|
||||
//! in-process (pure Rust, NUMA-safe). This crate is the **fallback** path:
|
||||
//! it offloads embedding to a local/remote Ollama daemon (`snowflake-arctic-embed2`),
|
||||
//! which is exactly the route the recall measurements used — so it reproduces
|
||||
//! the measured numbers (recall@10 130/132) byte-for-route. Opt-in via
|
||||
//! This crate offloads embedding to a local/remote Ollama daemon
|
||||
//! (`snowflake-arctic-embed2`), which is exactly the route the recall
|
||||
//! measurements used — so it reproduces the measured numbers (recall@10
|
||||
//! 130/132) byte-for-route. Opt-in via
|
||||
//! `config.models.embedding.provider = "ollama"`.
|
||||
//!
|
||||
//! ## Wire shape
|
||||
@@ -44,6 +43,7 @@
|
||||
use std::time::Duration;
|
||||
|
||||
use anyhow::{Context, Result};
|
||||
use kebab_config::EmbeddingModelCfg;
|
||||
use kebab_core::{Embedder, EmbeddingInput, EmbeddingKind, EmbeddingModelId, EmbeddingVersion};
|
||||
use serde::{Deserialize, Serialize};
|
||||
|
||||
@@ -64,10 +64,9 @@ const REQUEST_TIMEOUT_SECS: u64 = 300;
|
||||
|
||||
/// Resolve the (query_prefix, doc_prefix) for an Ollama embedding model tag.
|
||||
///
|
||||
/// Mirrors `kebab-embed-candle`'s `MODEL_REGISTRY`, but keyed on the **Ollama
|
||||
/// model tag** (which differs from the HF id — e.g. `snowflake-arctic-embed2`
|
||||
/// vs `Snowflake/snowflake-arctic-embed-l-v2.0`). Kept here rather than shared
|
||||
/// so this crate does not depend on the candle backend.
|
||||
/// Resolve the (query_prefix, doc_prefix) for an Ollama embedding model tag,
|
||||
/// keyed on the **Ollama model tag** (which differs from the HF id — e.g.
|
||||
/// `snowflake-arctic-embed2` vs `Snowflake/snowflake-arctic-embed-l-v2.0`).
|
||||
///
|
||||
/// An unrecognized model gets no prefix (`("", "")`): many embedding models
|
||||
/// are not instruction-tuned, so embedding the raw text is the correct default
|
||||
@@ -103,19 +102,15 @@ pub struct OllamaEmbedder {
|
||||
}
|
||||
|
||||
impl OllamaEmbedder {
|
||||
/// Build from a workspace [`kebab_config::Config`]. Reads
|
||||
/// `config.models.embedding.{model, dimensions}` and resolves the endpoint
|
||||
/// as `models.embedding.endpoint` → fallback `models.llm.endpoint`.
|
||||
/// Build from the `[models.embedding]` slice + a resolved `endpoint`.
|
||||
/// Reads `cfg.{model, dimensions}`; the caller resolves the endpoint
|
||||
/// (`models.embedding.endpoint` → fallback `models.llm.endpoint`) and
|
||||
/// passes it in.
|
||||
///
|
||||
/// Does NOT touch the network. The caller (app layer) is expected to have
|
||||
/// validated `provider == "ollama"`.
|
||||
pub fn new(config: &kebab_config::Config) -> Result<Self> {
|
||||
let emb = &config.models.embedding;
|
||||
let endpoint = emb
|
||||
.endpoint
|
||||
.clone()
|
||||
.filter(|e| !e.is_empty())
|
||||
.unwrap_or_else(|| config.models.llm.endpoint.clone());
|
||||
pub fn new(cfg: &EmbeddingModelCfg, endpoint: String) -> Result<Self> {
|
||||
let emb = cfg;
|
||||
if endpoint.is_empty() {
|
||||
anyhow::bail!(
|
||||
"ollama embedding provider needs an endpoint: set \
|
||||
|
||||
@@ -27,7 +27,16 @@ async fn embed_blocking(
|
||||
inputs: Vec<(String, EmbeddingKind)>,
|
||||
) -> anyhow::Result<Vec<Vec<f32>>> {
|
||||
tokio::task::spawn_blocking(move || -> anyhow::Result<Vec<Vec<f32>>> {
|
||||
let emb = OllamaEmbedder::new(&cfg)?;
|
||||
// Resolve the endpoint exactly as kebab-app's `embedder()` does:
|
||||
// `models.embedding.endpoint` → fallback `models.llm.endpoint`.
|
||||
let endpoint = cfg
|
||||
.models
|
||||
.embedding
|
||||
.endpoint
|
||||
.clone()
|
||||
.filter(|e| !e.is_empty())
|
||||
.unwrap_or_else(|| cfg.models.llm.endpoint.clone());
|
||||
let emb = OllamaEmbedder::new(&cfg.models.embedding, endpoint)?;
|
||||
let refs: Vec<EmbeddingInput<'_>> = inputs
|
||||
.iter()
|
||||
.map(|(t, k)| EmbeddingInput { text: t, kind: *k })
|
||||
|
||||
@@ -91,7 +91,7 @@ pub fn compare_runs_with_config(
|
||||
run_id_b: &str,
|
||||
opts: &CompareOpts,
|
||||
) -> Result<CompareReport> {
|
||||
let store = SqliteStore::open(cfg).context("open SqliteStore for compare_runs")?;
|
||||
let store = SqliteStore::open(&cfg.storage).context("open SqliteStore for compare_runs")?;
|
||||
store.run_migrations().context("run migrations")?;
|
||||
|
||||
// Pull both run rows up-front so we can extract chunker_version and
|
||||
|
||||
@@ -127,7 +127,7 @@ pub(crate) fn validate_against_db(
|
||||
return Ok(());
|
||||
}
|
||||
|
||||
let store = SqliteStore::open(cfg).context("open SqliteStore for golden validation")?;
|
||||
let store = SqliteStore::open(&cfg.storage).context("open SqliteStore for golden validation")?;
|
||||
store
|
||||
.run_migrations()
|
||||
.context("run migrations for golden validation")?;
|
||||
@@ -232,7 +232,7 @@ mod tests {
|
||||
let mut config = Config::defaults();
|
||||
config.storage.data_dir = tmp.path().to_string_lossy().into_owned();
|
||||
|
||||
let store = SqliteStore::open(&config).unwrap();
|
||||
let store = SqliteStore::open(&config.storage).unwrap();
|
||||
store.run_migrations().unwrap();
|
||||
seed_one_chunk(&store, "doc_present", "chunk_present");
|
||||
|
||||
@@ -256,7 +256,7 @@ mod tests {
|
||||
let mut config = Config::defaults();
|
||||
config.storage.data_dir = tmp.path().to_string_lossy().into_owned();
|
||||
|
||||
let store = SqliteStore::open(&config).unwrap();
|
||||
let store = SqliteStore::open(&config.storage).unwrap();
|
||||
store.run_migrations().unwrap();
|
||||
seed_one_chunk(&store, "doc_present", "chunk_present");
|
||||
|
||||
|
||||
@@ -114,7 +114,7 @@ pub fn compute_aggregate(run_id: &str) -> Result<AggregateMetrics> {
|
||||
/// Compute aggregate metrics for `run_id` against an explicit
|
||||
/// [`Config`] (used by tests with a TempDir-backed `data_dir`).
|
||||
pub fn compute_aggregate_with_config(cfg: &Config, run_id: &str) -> Result<AggregateMetrics> {
|
||||
let store = SqliteStore::open(cfg).context("open SqliteStore for compute_aggregate")?;
|
||||
let store = SqliteStore::open(&cfg.storage).context("open SqliteStore for compute_aggregate")?;
|
||||
store
|
||||
.run_migrations()
|
||||
.context("run migrations for compute_aggregate")?;
|
||||
@@ -146,7 +146,7 @@ pub fn store_aggregate_with_config(
|
||||
run_id: &str,
|
||||
agg: &AggregateMetrics,
|
||||
) -> Result<()> {
|
||||
let store = SqliteStore::open(cfg).context("open SqliteStore for store_aggregate")?;
|
||||
let store = SqliteStore::open(&cfg.storage).context("open SqliteStore for store_aggregate")?;
|
||||
store.run_migrations().context("run migrations")?;
|
||||
let json = serde_json::to_string(agg).context("serialize AggregateMetrics")?;
|
||||
store
|
||||
@@ -569,8 +569,6 @@ mod tests {
|
||||
latency_ms: 1,
|
||||
},
|
||||
created_at: OffsetDateTime::UNIX_EPOCH,
|
||||
conversation_id: None,
|
||||
turn_index: None,
|
||||
hops: None,
|
||||
verification: None,
|
||||
}
|
||||
|
||||
@@ -61,7 +61,7 @@ pub fn run_eval_with_config(cfg: &kebab_config::Config, opts: &EvalRunOpts) -> R
|
||||
|
||||
// Open the store once so every per-query write reuses the same
|
||||
// connection-mutex lifetime.
|
||||
let store = SqliteStore::open(cfg).context("open SqliteStore for run_eval")?;
|
||||
let store = SqliteStore::open(&cfg.storage).context("open SqliteStore for run_eval")?;
|
||||
store
|
||||
.run_migrations()
|
||||
.context("run migrations for run_eval")?;
|
||||
@@ -174,11 +174,6 @@ fn execute_query(app: &App, gq: &GoldenQuery, opts: &EvalRunOpts) -> QueryResult
|
||||
temperature: opts.temperature,
|
||||
seed: opts.seed,
|
||||
stream_sink: None,
|
||||
// p9-fb-15: golden eval is single-shot per query; no
|
||||
// conversational history.
|
||||
history: Vec::new(),
|
||||
conversation_id: None,
|
||||
turn_index: None,
|
||||
// p9-fb-41: golden eval baseline runs are single-pass; the
|
||||
// multi-hop path is opted into per query via a future
|
||||
// fixture flag (PR-4+) once the runner learns to dispatch.
|
||||
|
||||
@@ -239,7 +239,7 @@ pub fn compute_variant_consistency_with_config(
|
||||
cfg: &Config,
|
||||
run_id: &str,
|
||||
) -> Result<VariantConsistencyReport> {
|
||||
let store = SqliteStore::open(cfg).context("open SqliteStore for variant consistency")?;
|
||||
let store = SqliteStore::open(&cfg.storage).context("open SqliteStore for variant consistency")?;
|
||||
store.run_migrations().context("run migrations")?;
|
||||
let run_record = store
|
||||
.load_eval_run(run_id)
|
||||
|
||||
@@ -152,7 +152,7 @@ fn compute_and_store_aggregate_round_trips() {
|
||||
let _g = env_guard();
|
||||
let tmp = TempDir::new().unwrap();
|
||||
let cfg = cfg_with_data_dir(&tmp, golden_yaml_basic());
|
||||
let store = SqliteStore::open(&cfg).unwrap();
|
||||
let store = SqliteStore::open(&cfg.storage).unwrap();
|
||||
store.run_migrations().unwrap();
|
||||
let now = OffsetDateTime::UNIX_EPOCH;
|
||||
write_run(
|
||||
@@ -183,7 +183,7 @@ fn compute_and_store_aggregate_round_trips() {
|
||||
assert_eq!(agg.mrr, 0.4167);
|
||||
|
||||
store_aggregate_with_config(&cfg, "run_a", &agg).unwrap();
|
||||
let store = SqliteStore::open(&cfg).unwrap();
|
||||
let store = SqliteStore::open(&cfg.storage).unwrap();
|
||||
let row = store.load_eval_run("run_a").unwrap().unwrap();
|
||||
let parsed: AggregateMetrics = serde_json::from_str(&row.aggregate_json).unwrap();
|
||||
// f32 round-trip via JSON is exact for our 4-decimal-rounded
|
||||
@@ -224,7 +224,7 @@ fn compare_runs_classifies_win_loss_draw_regression() {
|
||||
let _g = env_guard();
|
||||
let tmp = TempDir::new().unwrap();
|
||||
let cfg = cfg_with_data_dir(&tmp, golden_yaml_basic());
|
||||
let store = SqliteStore::open(&cfg).unwrap();
|
||||
let store = SqliteStore::open(&cfg.storage).unwrap();
|
||||
store.run_migrations().unwrap();
|
||||
let now = OffsetDateTime::UNIX_EPOCH;
|
||||
// Run A:
|
||||
@@ -284,7 +284,7 @@ fn compare_strict_mode_refuses_chunker_version_mismatch() {
|
||||
let _g = env_guard();
|
||||
let tmp = TempDir::new().unwrap();
|
||||
let cfg = cfg_with_data_dir(&tmp, golden_yaml_basic());
|
||||
let store = SqliteStore::open(&cfg).unwrap();
|
||||
let store = SqliteStore::open(&cfg.storage).unwrap();
|
||||
store.run_migrations().unwrap();
|
||||
let now = OffsetDateTime::UNIX_EPOCH;
|
||||
write_run(
|
||||
@@ -316,7 +316,7 @@ fn compare_graceful_falls_back_to_doc_id() {
|
||||
let _g = env_guard();
|
||||
let tmp = TempDir::new().unwrap();
|
||||
let cfg = cfg_with_data_dir(&tmp, golden_yaml_basic());
|
||||
let store = SqliteStore::open(&cfg).unwrap();
|
||||
let store = SqliteStore::open(&cfg.storage).unwrap();
|
||||
store.run_migrations().unwrap();
|
||||
let now = OffsetDateTime::UNIX_EPOCH;
|
||||
// Run A uses test@1 chunker; run B uses test@2 — chunk_ids no longer
|
||||
@@ -357,7 +357,7 @@ fn compare_report_snapshot_matches_fixture() {
|
||||
let _g = env_guard();
|
||||
let tmp = TempDir::new().unwrap();
|
||||
let cfg = cfg_with_data_dir(&tmp, golden_yaml_basic());
|
||||
let store = SqliteStore::open(&cfg).unwrap();
|
||||
let store = SqliteStore::open(&cfg.storage).unwrap();
|
||||
store.run_migrations().unwrap();
|
||||
let now = OffsetDateTime::UNIX_EPOCH;
|
||||
write_run(
|
||||
@@ -434,7 +434,7 @@ fn render_report_md_is_human_readable() {
|
||||
let _g = env_guard();
|
||||
let tmp = TempDir::new().unwrap();
|
||||
let cfg = cfg_with_data_dir(&tmp, golden_yaml_basic());
|
||||
let store = SqliteStore::open(&cfg).unwrap();
|
||||
let store = SqliteStore::open(&cfg.storage).unwrap();
|
||||
store.run_migrations().unwrap();
|
||||
let now = OffsetDateTime::UNIX_EPOCH;
|
||||
write_run(
|
||||
|
||||
@@ -48,7 +48,7 @@ impl RunEnv {
|
||||
// Pin search defaults so test asserts are stable.
|
||||
config.search.default_k = 5;
|
||||
|
||||
let store = SqliteStore::open(&config).unwrap();
|
||||
let store = SqliteStore::open(&config.storage).unwrap();
|
||||
store.run_migrations().unwrap();
|
||||
seed_corpus(&store);
|
||||
Self { temp, config }
|
||||
@@ -273,7 +273,7 @@ fn runner_persists_eval_run_and_query_result_rows() {
|
||||
// the rows back. We use the inherent `read_conn` helper rather
|
||||
// than rusqlite directly because the latter would require kb-eval
|
||||
// to add a runtime rusqlite dep (forbidden by the spec).
|
||||
let store = SqliteStore::open(&env.config).unwrap();
|
||||
let store = SqliteStore::open(&env.config.storage).unwrap();
|
||||
let conn = store.read_conn();
|
||||
|
||||
let n_runs: i64 = conn
|
||||
|
||||
@@ -1,6 +1,5 @@
|
||||
//! `ask` tool — wraps `kebab_app::ask_with_config` (single-shot) or
|
||||
//! `kebab_app::ask_with_session_with_config` when `session_id` is provided.
|
||||
//! Input: { query, session_id?, mode? }. Output: answer.v1 JSON.
|
||||
//! `ask` tool — wraps `kebab_app::ask_with_config` (single-shot).
|
||||
//! Input: { query, mode? }. Output: answer.v1 JSON.
|
||||
//!
|
||||
//! `Answer` (kebab-core) does NOT carry a `schema_version` field; we tag
|
||||
//! it inline here, matching the pattern from `search.rs`.
|
||||
@@ -16,8 +15,6 @@ use crate::state::KebabAppState;
|
||||
pub struct AskInput {
|
||||
/// The user question.
|
||||
pub query: String,
|
||||
/// Optional session id for multi-turn RAG context.
|
||||
pub session_id: Option<String>,
|
||||
/// Optional retrieval mode override ("lexical" / "vector" / "hybrid"). Default "hybrid".
|
||||
pub mode: Option<String>,
|
||||
/// p9-fb-41: opt the ask into the multi-hop pipeline. Default `false`.
|
||||
@@ -44,16 +41,10 @@ pub fn handle(state: &KebabAppState, input: AskInput) -> CallToolResult {
|
||||
temperature: None,
|
||||
seed: None,
|
||||
stream_sink: None,
|
||||
history: Vec::new(),
|
||||
conversation_id: None,
|
||||
turn_index: None,
|
||||
multi_hop: input.multi_hop.unwrap_or(false),
|
||||
};
|
||||
let cfg_clone = (*state.config).clone();
|
||||
let result = match input.session_id {
|
||||
Some(sid) => kebab_app::ask_with_session_with_config(cfg_clone, &sid, &input.query, opts),
|
||||
None => kebab_app::ask_with_config(cfg_clone, &input.query, opts),
|
||||
};
|
||||
let result = kebab_app::ask_with_config(cfg_clone, &input.query, opts);
|
||||
match result {
|
||||
Ok(answer) => {
|
||||
// `Answer` does not carry `schema_version`; tag inline (idempotent
|
||||
|
||||
@@ -33,7 +33,7 @@ async fn ask_tool_returns_answer_v1_with_refusal_on_empty_kb() {
|
||||
include: vec![],
|
||||
exclude: vec![],
|
||||
};
|
||||
let _ = kebab_app::ingest_with_config(cfg.clone(), scope, false).unwrap();
|
||||
let _ = kebab_app::ingest_with_config(cfg.clone(), scope, kebab_app::IngestOpts::default()).unwrap();
|
||||
|
||||
let state = KebabAppState::new(cfg, None);
|
||||
let handler = KebabHandler::new(state);
|
||||
@@ -48,7 +48,6 @@ async fn ask_tool_returns_answer_v1_with_refusal_on_empty_kb() {
|
||||
&state_clone,
|
||||
kebab_mcp::tools::ask::AskInput {
|
||||
query: "what is the meaning of life".to_string(),
|
||||
session_id: None,
|
||||
// Test env uses provider="none" — Hybrid would hard-error on embedding.
|
||||
// Pass Lexical explicitly so the test stays functional.
|
||||
mode: Some("lexical".to_string()),
|
||||
|
||||
@@ -90,7 +90,7 @@ async fn ask_tool_routes_multi_hop_true_to_decompose_first() {
|
||||
include: vec![],
|
||||
exclude: vec![],
|
||||
};
|
||||
let _ = kebab_app::ingest_with_config(cfg.clone(), scope, false).unwrap();
|
||||
let _ = kebab_app::ingest_with_config(cfg.clone(), scope, kebab_app::IngestOpts::default()).unwrap();
|
||||
|
||||
let state = KebabAppState::new(cfg, None);
|
||||
let handler = KebabHandler::new(state);
|
||||
@@ -103,7 +103,6 @@ async fn ask_tool_routes_multi_hop_true_to_decompose_first() {
|
||||
&state_mh,
|
||||
kebab_mcp::tools::ask::AskInput {
|
||||
query: "compound about X and Y".to_string(),
|
||||
session_id: None,
|
||||
mode: Some("lexical".to_string()),
|
||||
multi_hop: Some(true),
|
||||
},
|
||||
@@ -138,7 +137,6 @@ async fn ask_tool_routes_multi_hop_true_to_decompose_first() {
|
||||
&state_sp,
|
||||
kebab_mcp::tools::ask::AskInput {
|
||||
query: "anything".to_string(),
|
||||
session_id: None,
|
||||
mode: Some("lexical".to_string()),
|
||||
multi_hop: Some(false),
|
||||
},
|
||||
@@ -179,7 +177,7 @@ async fn ask_tool_multi_hop_short_circuits_when_probe_empty() {
|
||||
include: vec![],
|
||||
exclude: vec![],
|
||||
};
|
||||
let _ = kebab_app::ingest_with_config(cfg.clone(), scope, false).unwrap();
|
||||
let _ = kebab_app::ingest_with_config(cfg.clone(), scope, kebab_app::IngestOpts::default()).unwrap();
|
||||
|
||||
let state = KebabAppState::new(cfg.clone(), None);
|
||||
let handler = KebabHandler::new(state);
|
||||
@@ -189,7 +187,6 @@ async fn ask_tool_multi_hop_short_circuits_when_probe_empty() {
|
||||
&state_mh,
|
||||
kebab_mcp::tools::ask::AskInput {
|
||||
query: "compound about X and Y".to_string(),
|
||||
session_id: None,
|
||||
mode: Some("lexical".to_string()),
|
||||
multi_hop: Some(true),
|
||||
},
|
||||
|
||||
@@ -36,7 +36,7 @@ fn setup() -> (tempfile::TempDir, KebabHandler) {
|
||||
include: vec![],
|
||||
exclude: vec![],
|
||||
};
|
||||
let _ = kebab_app::ingest_with_config(config.clone(), scope, false).unwrap();
|
||||
let _ = kebab_app::ingest_with_config(config.clone(), scope, kebab_app::IngestOpts::default()).unwrap();
|
||||
let state = KebabAppState::new(config, None);
|
||||
let handler = KebabHandler::new(state);
|
||||
(dir, handler)
|
||||
|
||||
@@ -44,7 +44,7 @@ async fn fetch_tool_chunk_returns_fetch_result_v1() {
|
||||
include: vec![],
|
||||
exclude: vec![],
|
||||
};
|
||||
let _ = kebab_app::ingest_with_config(config.clone(), scope, false).unwrap();
|
||||
let _ = kebab_app::ingest_with_config(config.clone(), scope, kebab_app::IngestOpts::default()).unwrap();
|
||||
|
||||
let state = KebabAppState::new(config, None);
|
||||
let handler = KebabHandler::new(state);
|
||||
|
||||
@@ -34,7 +34,7 @@ async fn schema_tool_returns_schema_v1_json() {
|
||||
include: vec![],
|
||||
exclude: vec![],
|
||||
};
|
||||
let _ = kebab_app::ingest_with_config(config.clone(), scope, false).unwrap();
|
||||
let _ = kebab_app::ingest_with_config(config.clone(), scope, kebab_app::IngestOpts::default()).unwrap();
|
||||
|
||||
let state = KebabAppState::new(config, None);
|
||||
let handler = KebabHandler::new(state);
|
||||
|
||||
@@ -41,7 +41,7 @@ async fn search_tool_returns_search_response_v1() {
|
||||
include: vec![],
|
||||
exclude: vec![],
|
||||
};
|
||||
let _ = kebab_app::ingest_with_config(config.clone(), scope, false).unwrap();
|
||||
let _ = kebab_app::ingest_with_config(config.clone(), scope, kebab_app::IngestOpts::default()).unwrap();
|
||||
|
||||
let state = KebabAppState::new(config, None);
|
||||
let handler = KebabHandler::new(state);
|
||||
@@ -141,7 +141,7 @@ async fn search_with_doc_id_filter_returns_only_target() {
|
||||
include: vec![],
|
||||
exclude: vec![],
|
||||
};
|
||||
let _ = kebab_app::ingest_with_config(config.clone(), scope, false).unwrap();
|
||||
let _ = kebab_app::ingest_with_config(config.clone(), scope, kebab_app::IngestOpts::default()).unwrap();
|
||||
|
||||
let state = KebabAppState::new(config, None);
|
||||
let handler = KebabHandler::new(state);
|
||||
|
||||
@@ -35,7 +35,7 @@ fn setup() -> (tempfile::TempDir, KebabHandler) {
|
||||
include: vec![],
|
||||
exclude: vec![],
|
||||
};
|
||||
let _ = kebab_app::ingest_with_config(config.clone(), scope, false).unwrap();
|
||||
let _ = kebab_app::ingest_with_config(config.clone(), scope, kebab_app::IngestOpts::default()).unwrap();
|
||||
let state = KebabAppState::new(config, None);
|
||||
let handler = KebabHandler::new(state);
|
||||
(dir, handler)
|
||||
|
||||
@@ -438,6 +438,8 @@ pub(crate) mod tests_support {
|
||||
asset: &asset,
|
||||
workspace_root: &root,
|
||||
config: &cfg,
|
||||
source_id: None,
|
||||
source_trust: None,
|
||||
};
|
||||
CAstExtractor::new().extract(&ctx, src.as_bytes()).unwrap()
|
||||
}
|
||||
|
||||
@@ -646,6 +646,8 @@ pub(crate) mod tests_support {
|
||||
asset: &asset,
|
||||
workspace_root: &root,
|
||||
config: &cfg,
|
||||
source_id: None,
|
||||
source_trust: None,
|
||||
};
|
||||
CppAstExtractor::new()
|
||||
.extract(&ctx, src.as_bytes())
|
||||
|
||||
@@ -392,6 +392,8 @@ mod tests {
|
||||
asset: &asset,
|
||||
workspace_root: &root,
|
||||
config: &cfg,
|
||||
source_id: None,
|
||||
source_trust: None,
|
||||
};
|
||||
GoAstExtractor::new().extract(&ctx, &bytes).unwrap()
|
||||
}
|
||||
|
||||
@@ -454,6 +454,8 @@ mod tests {
|
||||
asset: &asset,
|
||||
workspace_root: &root,
|
||||
config: &cfg,
|
||||
source_id: None,
|
||||
source_trust: None,
|
||||
};
|
||||
JavaAstExtractor::new().extract(&ctx, &bytes).unwrap()
|
||||
}
|
||||
|
||||
@@ -460,6 +460,8 @@ mod tests {
|
||||
asset: &asset,
|
||||
workspace_root: &root,
|
||||
config: &cfg,
|
||||
source_id: None,
|
||||
source_trust: None,
|
||||
};
|
||||
JavascriptAstExtractor::new().extract(&ctx, &bytes).unwrap()
|
||||
}
|
||||
@@ -510,6 +512,8 @@ mod tests {
|
||||
asset: &asset,
|
||||
workspace_root: &root,
|
||||
config: &cfg,
|
||||
source_id: None,
|
||||
source_trust: None,
|
||||
};
|
||||
let doc = JavascriptAstExtractor::new().extract(&ctx, bytes).unwrap();
|
||||
let syms = symbols(&doc);
|
||||
@@ -542,6 +546,8 @@ mod tests {
|
||||
asset: &asset,
|
||||
workspace_root: &root,
|
||||
config: &cfg,
|
||||
source_id: None,
|
||||
source_trust: None,
|
||||
};
|
||||
let doc = JavascriptAstExtractor::new().extract(&ctx, bytes).unwrap();
|
||||
|
||||
|
||||
@@ -532,6 +532,8 @@ mod tests {
|
||||
asset: &asset,
|
||||
workspace_root: &root,
|
||||
config: &cfg,
|
||||
source_id: None,
|
||||
source_trust: None,
|
||||
};
|
||||
KotlinAstExtractor::new().extract(&ctx, &bytes).unwrap()
|
||||
}
|
||||
|
||||
@@ -397,6 +397,8 @@ mod tests {
|
||||
asset: &asset,
|
||||
workspace_root: &root,
|
||||
config: &cfg,
|
||||
source_id: None,
|
||||
source_trust: None,
|
||||
};
|
||||
PythonAstExtractor::new().extract(&ctx, &bytes).unwrap()
|
||||
}
|
||||
|
||||
@@ -400,6 +400,8 @@ mod tests {
|
||||
asset: &asset,
|
||||
workspace_root: &root,
|
||||
config: &cfg,
|
||||
source_id: None,
|
||||
source_trust: None,
|
||||
};
|
||||
RustAstExtractor::new().extract(&ctx, &bytes).unwrap()
|
||||
}
|
||||
@@ -463,6 +465,8 @@ mod tests {
|
||||
asset: &asset,
|
||||
workspace_root: &root,
|
||||
config: &cfg,
|
||||
source_id: None,
|
||||
source_trust: None,
|
||||
};
|
||||
let doc = RustAstExtractor::new()
|
||||
.extract(&ctx, source.as_bytes())
|
||||
|
||||
@@ -501,6 +501,8 @@ mod tests {
|
||||
asset: &asset,
|
||||
workspace_root: &root,
|
||||
config: &cfg,
|
||||
source_id: None,
|
||||
source_trust: None,
|
||||
};
|
||||
TypescriptAstExtractor::new().extract(&ctx, &bytes).unwrap()
|
||||
}
|
||||
@@ -587,6 +589,8 @@ mod tests {
|
||||
asset: &asset,
|
||||
workspace_root: &root,
|
||||
config: &cfg,
|
||||
source_id: None,
|
||||
source_trust: None,
|
||||
};
|
||||
let doc = TypescriptAstExtractor::new().extract(&ctx, bytes).unwrap();
|
||||
|
||||
@@ -640,6 +644,8 @@ mod tests {
|
||||
asset: &asset,
|
||||
workspace_root: &root,
|
||||
config: &cfg,
|
||||
source_id: None,
|
||||
source_trust: None,
|
||||
};
|
||||
let doc = TypescriptAstExtractor::new().extract(&ctx, bytes).unwrap();
|
||||
|
||||
|
||||
@@ -133,7 +133,7 @@ impl OllamaVisionOcr {
|
||||
/// Construction does NOT touch the network — the first HTTP call
|
||||
/// happens inside [`OcrEngine::recognize`].
|
||||
pub fn new(config: &kebab_config::Config) -> Result<Self> {
|
||||
let ocr = &config.ingest.image.ocr;
|
||||
let ocr = config.image_ocr();
|
||||
let endpoint = match ocr.endpoint.as_deref() {
|
||||
Some(s) if !s.is_empty() => s.to_string(),
|
||||
_ => config.models.llm.endpoint.clone(),
|
||||
|
||||
@@ -122,7 +122,7 @@ impl ModelPaths {
|
||||
/// [`from_default_dir`]: ModelPaths::from_default_dir
|
||||
pub fn from_config(config: &kebab_config::Config) -> Self {
|
||||
let defaults = Self::from_default_dir();
|
||||
let ocr = &config.ingest.image.ocr;
|
||||
let ocr = config.image_ocr();
|
||||
Self {
|
||||
det: ocr.det_model.as_ref().map(PathBuf::from).unwrap_or(defaults.det),
|
||||
rec: ocr.rec_model.as_ref().map(PathBuf::from).unwrap_or(defaults.rec),
|
||||
@@ -138,7 +138,7 @@ impl OnnxPaddleOcr {
|
||||
/// here are fail-fast (matches the Ollama adapter's construction contract).
|
||||
pub fn new(config: &kebab_config::Config) -> Result<Self> {
|
||||
let paths = ModelPaths::from_config(config);
|
||||
let ocr = &config.ingest.image.ocr;
|
||||
let ocr = config.image_ocr();
|
||||
Self::from_paths(
|
||||
&paths,
|
||||
ocr.score_thresh,
|
||||
|
||||
@@ -282,6 +282,8 @@ impl ImageFixture {
|
||||
asset: &self.asset,
|
||||
workspace_root: &self.workspace_root,
|
||||
config: &self.config,
|
||||
source_id: None,
|
||||
source_trust: None,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -354,15 +354,16 @@ fn from_parts_clamps_max_pixels_into_legal_range() {
|
||||
/// Run with:
|
||||
///
|
||||
/// ```sh
|
||||
/// KEBAB_IMAGE_OCR_ENDPOINT=http://192.168.0.47:11434 \
|
||||
/// KEBAB_OCR_ENDPOINT=http://192.168.0.47:11434 \
|
||||
/// cargo test -p kebab-parse-image --test ocr ocr_integration -- --ignored
|
||||
/// ```
|
||||
#[tokio::test]
|
||||
#[ignore = "hits a real Ollama daemon; opt in via `cargo test -- --ignored`"]
|
||||
async fn ocr_integration_real_ollama_transcribes_text() {
|
||||
let endpoint = std::env::var("KEBAB_IMAGE_OCR_ENDPOINT")
|
||||
// v5: shared KEBAB_OCR_* env (manual harness reads it directly).
|
||||
let endpoint = std::env::var("KEBAB_OCR_ENDPOINT")
|
||||
.unwrap_or_else(|_| "http://192.168.0.47:11434".to_string());
|
||||
let model = std::env::var("KEBAB_IMAGE_OCR_MODEL").unwrap_or_else(|_| "gemma4:e4b".to_string());
|
||||
let model = std::env::var("KEBAB_OCR_MODEL").unwrap_or_else(|_| "gemma4:e4b".to_string());
|
||||
|
||||
// Generate a fixture with known text. If the DejaVu font is
|
||||
// missing from this dev box, skip rather than crash.
|
||||
|
||||
122
crates/kebab-parse-md/src/extractor.rs
Normal file
122
crates/kebab-parse-md/src/extractor.rs
Normal file
@@ -0,0 +1,122 @@
|
||||
//! `kb-parse-md::extractor` — the [`Extractor`] trait impl that wraps the
|
||||
//! crate's free functions (`parse_frontmatter` + `parse_blocks` +
|
||||
//! `build_canonical_document`) so markdown ingest flows through the same
|
||||
//! `App.extractors` registry + `App::extract_for` polymorphic dispatch
|
||||
//! that pdf / image / code already use.
|
||||
//!
|
||||
//! This is a pure structural unification: the byte sequence it runs is
|
||||
//! identical to the inline arm `kebab-app::ingest_one_asset` used before
|
||||
//! (frontmatter parse → body-offset count → block parse → canonical lift,
|
||||
//! same args, same order), so the produced `CanonicalDocument` — and thus
|
||||
//! `doc_id` / `chunk_id` — is byte-for-byte the same.
|
||||
//!
|
||||
//! The one piece of context the inline arm read that the other extractors
|
||||
//! do not is the per-source `source_id` + `trust_level`: markdown
|
||||
//! frontmatter can *override* the per-source trust default, and that
|
||||
//! precedence is resolved *inside* `parse_frontmatter` via [`BodyHints`].
|
||||
//! [`ExtractContext`] carries both so the resolution stays identical.
|
||||
|
||||
use kebab_core::{CanonicalDocument, ExtractContext, Extractor, MediaType, ParserVersion, RawAsset};
|
||||
|
||||
use crate::PARSER_VERSION;
|
||||
use crate::frontmatter::{BodyHints, FrontmatterSpan, parse_frontmatter};
|
||||
use crate::{build_canonical_document, parse_blocks};
|
||||
|
||||
/// Markdown extractor — wraps the crate's free functions behind the
|
||||
/// [`Extractor`] trait.
|
||||
pub struct MarkdownExtractor;
|
||||
|
||||
impl MarkdownExtractor {
|
||||
pub fn new() -> Self {
|
||||
Self
|
||||
}
|
||||
}
|
||||
|
||||
impl Default for MarkdownExtractor {
|
||||
fn default() -> Self {
|
||||
Self::new()
|
||||
}
|
||||
}
|
||||
|
||||
impl Extractor for MarkdownExtractor {
|
||||
fn supports(&self, m: &MediaType) -> bool {
|
||||
matches!(m, MediaType::Markdown)
|
||||
}
|
||||
|
||||
fn parser_version(&self) -> ParserVersion {
|
||||
ParserVersion(PARSER_VERSION.to_string())
|
||||
}
|
||||
|
||||
fn extract(
|
||||
&self,
|
||||
ctx: &ExtractContext<'_>,
|
||||
bytes: &[u8],
|
||||
) -> anyhow::Result<CanonicalDocument> {
|
||||
let asset = ctx.asset;
|
||||
let parser_version = self.parser_version();
|
||||
|
||||
// `[[workspace.sources]]`: stamp the owning source id + inject the
|
||||
// per-source default trust level (frontmatter still overrides it).
|
||||
// Mirrors the old inline `build_body_hints` exactly.
|
||||
let body_hints = build_body_hints(asset, ctx.source_id, ctx.source_trust);
|
||||
|
||||
// Frontmatter — `parse_frontmatter` returns Ok even on malformed
|
||||
// frontmatter (warnings are surfaced through the `Vec<Warning>`).
|
||||
use anyhow::Context as _;
|
||||
let (metadata, fm_span, fm_warns) =
|
||||
parse_frontmatter(bytes, &body_hints).context("kb-parse-md::parse_frontmatter")?;
|
||||
|
||||
let body_offset_lines = match fm_span {
|
||||
Some(span) => count_lines_in(&bytes[..span.end]),
|
||||
None => 0,
|
||||
};
|
||||
|
||||
let (parsed_blocks, blk_warns) =
|
||||
parse_blocks(&bytes[fm_span_end(fm_span)..], body_offset_lines)
|
||||
.context("kb-parse-md::parse_blocks")?;
|
||||
|
||||
let mut all_warnings = Vec::with_capacity(fm_warns.len() + blk_warns.len());
|
||||
all_warnings.extend(fm_warns);
|
||||
all_warnings.extend(blk_warns);
|
||||
|
||||
let canonical =
|
||||
build_canonical_document(asset, metadata, parsed_blocks, &parser_version, all_warnings)
|
||||
.context("kb-parse-md::build_canonical_document")?;
|
||||
Ok(canonical)
|
||||
}
|
||||
}
|
||||
|
||||
/// Build `BodyHints` from the asset alone. We use the asset's
|
||||
/// `discovered_at` for both `fs_ctime` and `fs_mtime` because going
|
||||
/// through the FS metadata API for every file would be a noticeable
|
||||
/// overhead for large workspaces and the source-of-truth timestamps
|
||||
/// are written into the document's frontmatter when the user wants
|
||||
/// authoritative values.
|
||||
fn build_body_hints(
|
||||
asset: &RawAsset,
|
||||
source_id: Option<&str>,
|
||||
source_trust: Option<kebab_core::TrustLevel>,
|
||||
) -> BodyHints {
|
||||
BodyHints {
|
||||
first_h1: None,
|
||||
fs_ctime: asset.discovered_at,
|
||||
fs_mtime: asset.discovered_at,
|
||||
fallback_lang: None,
|
||||
// `[[workspace.sources]]`: stamp the owning source id + inject the
|
||||
// per-source default trust level (frontmatter still overrides it).
|
||||
source_id: source_id.map(str::to_string),
|
||||
fallback_trust_level: source_trust,
|
||||
}
|
||||
}
|
||||
|
||||
/// Convenience: end byte of the frontmatter region (or 0 when absent).
|
||||
fn fm_span_end(span: Option<FrontmatterSpan>) -> usize {
|
||||
span.map_or(0, |s| s.end)
|
||||
}
|
||||
|
||||
/// Count `\n` in a byte prefix to convert frontmatter byte span to
|
||||
/// the line-offset `parse_blocks` expects.
|
||||
fn count_lines_in(bytes: &[u8]) -> u32 {
|
||||
let n = bytes.iter().filter(|&&b| b == b'\n').count();
|
||||
u32::try_from(n).unwrap_or(u32::MAX)
|
||||
}
|
||||
@@ -18,6 +18,10 @@
|
||||
//! * [`build_canonical_document`] / [`derive_title`] — lift a parsed
|
||||
//! markdown document into a `kebab_core::CanonicalDocument` (absorbed
|
||||
//! from `kebab-normalize` — P1-4 / p9-fb-07 frozen API).
|
||||
//! * [`MarkdownExtractor`] — the [`kebab_core::Extractor`] impl that wraps
|
||||
//! the three free functions above so markdown ingest flows through the
|
||||
//! `App.extractors` registry like pdf / image / code (extract-stage
|
||||
//! symmetry).
|
||||
//! * Parser intermediate types ([`ParsedBlock`], [`ParsedBlockKind`],
|
||||
//! [`ParsedPayload`], [`Warning`], [`WarningKind`]) and 3 forward-declared
|
||||
//! structs ([`ParsedImageRegion`], [`ParsedPdfPage`], [`ParsedAudioSegment`]) —
|
||||
@@ -26,11 +30,13 @@
|
||||
//! Anything else in this crate is `pub(crate)` and may change without notice.
|
||||
|
||||
pub mod blocks;
|
||||
mod extractor;
|
||||
pub mod frontmatter;
|
||||
mod normalize;
|
||||
mod types;
|
||||
|
||||
pub use blocks::parse_blocks;
|
||||
pub use extractor::MarkdownExtractor;
|
||||
pub use frontmatter::{BodyHints, FrontmatterSpan, parse_frontmatter};
|
||||
|
||||
// Spec §3.3 의 surface 보존 정책 — explicit (NOT glob) 으로 future addition leak 방지.
|
||||
|
||||
@@ -169,6 +169,8 @@ impl PdfFixture {
|
||||
asset: &self.asset,
|
||||
workspace_root: &self.workspace_root,
|
||||
config: &self.config,
|
||||
source_id: None,
|
||||
source_trust: None,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
// F1 ≥ 0.85, F2 ≥ 0.70. real Ollama 의존 — `#[ignore]` default.
|
||||
//
|
||||
// Manual invoke:
|
||||
// KEBAB_PDF_OCR_ENDPOINT=http://192.168.0.47:11434 \
|
||||
// KEBAB_OCR_ENDPOINT=http://192.168.0.47:11434 \
|
||||
// cargo test -p kebab-parse-pdf --test ocr_e2e --ignored -j 4
|
||||
|
||||
use kebab_core::Lang;
|
||||
@@ -11,7 +11,8 @@ use kebab_parse_pdf::extract_dctdecode_page_image;
|
||||
use lopdf::Document;
|
||||
|
||||
fn run_real_ollama_ocr(pdf: &[u8], page: u32) -> anyhow::Result<String> {
|
||||
let endpoint = std::env::var("KEBAB_PDF_OCR_ENDPOINT")
|
||||
// v5: shared KEBAB_OCR_* env (manual harness reads it directly).
|
||||
let endpoint = std::env::var("KEBAB_OCR_ENDPOINT")
|
||||
.unwrap_or_else(|_| "http://localhost:11434".to_string());
|
||||
let doc = Document::load_mem(pdf)?;
|
||||
let jpeg = extract_dctdecode_page_image(&doc, page)?
|
||||
|
||||
@@ -47,6 +47,8 @@ fn vector_pdf_extract_byte_identical_to_baseline() {
|
||||
asset: &asset,
|
||||
workspace_root,
|
||||
config: &config,
|
||||
source_id: None,
|
||||
source_trust: None,
|
||||
};
|
||||
|
||||
let mut canonical = PdfTextExtractor::new()
|
||||
@@ -96,6 +98,8 @@ fn pdf_text_extractor_on_mojibake_yields_one_block() {
|
||||
asset: &asset,
|
||||
workspace_root,
|
||||
config: &config,
|
||||
source_id: None,
|
||||
source_trust: None,
|
||||
};
|
||||
let canonical = PdfTextExtractor::new()
|
||||
.extract(&ctx, bytes)
|
||||
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user