#239 수정(PR #240) 을 릴리스로 컷한다. 사용자 표면(서브커맨드·플래그·config 키·wire 스키마)에 변화가 없어 patch 로 판정했다. 규칙 문면상으로는 minor 로 읽히는 지점이 있으므로 근거를 Cargo.toml 주석과 릴리스 노트 앞머리에 남겼다. 릴리스 노트를 두 관점으로 사실검증하면서 **이미 머지된 HOTFIXES 항목의 사실 오류 두 개**가 드러나 함께 고쳤다. 1. "폭이 한 자릿수인 입력은 폭 0 인 특징맵이 된다" — 실제로는 폭 4 이하만 실패한다. 세 줄 아래 표(5~48 전부 성공)와 스스로 모순이었다. 이슈 본문의 표현을 그대로 물려받은 것이다. 2. "이미 인식된 87 개 영역을 통째로 잃고 있었다" — 87 은 수정 **후** 나오는 영역 수다. 계측으로 확인: Clickpath_Analysis.png 은 박스를 119 개 검출하고, 수정 전에는 그중 3 개를 인식한 시점에 중단됐다 (`ABORT regions_already_recognized=3 of detected_boxes=119`). 약 29 배 과장이었고, 하필 "이미 인식한 것까지 버려진다" 는 기제를 설명하는 자리라 숫자가 논지를 떠받치고 있었다. 릴리스 노트에서 추가로 바로잡은 것: - 컴파일 가드는 "올리면 실패" 가 아니라 "16 을 넘기면 실패" 다. 6~16 은 그대로 빌드되고 테스트도 통과한다. - "16 으로 하면 폭 5~15 를 새로 버린다" 는 근거가 정반대였다. 말뭉치 240 개 파일을 계측해 직접 확인: 세션에 투입된 박스 10,991 개 중 폭 5~16 은 403 개, 그중 글자를 뱉은 것은 0 개이고 글자가 처음 나오는 폭은 17 이다. 5 를 고른 이유는 손실이 아니라 상수 이름과 주석이 거짓이 되지 않기 때문이다. - "실패했던 문서만 다시 OCR 된다" 는 무조건문이 아니다. OCR 파생 캐시는 v0.31.0 에 들어왔는데 이미지 parser_version 은 v0.28.0 이후 안 바뀌었으므로, v0.31.0 이전에 색인하고 그 뒤 재처리된 적 없는 KB 는 캐시가 비어 성공했던 이미지까지 전부 다시 돈다. - 피해 집계를 절차 맨 뒤(5번)에 둬서, 순서대로 따르면 세기 전에 색인해 버리고 창이 영구히 닫혔다. 0 번으로 올리고 "반드시" 로 바꿨다. - `_external/` 로 넣은 문서는 `.kebabignore` 때문에 workspace 색인이 방문하지 않으므로 자동 재색인 대상이 아니다. - `pdf_ocr_events` 대안 쿼리는 30 일 보관 정리를 받고 `'ocr_error'` 가 모든 OCR 실패를 받는 통칭이라 상한이다. 0 이 나와도 안 당한 게 아니다. - doc_id 가 전부 바뀌면 저장해 둔 인용과 integrations/claude-code 스킬처럼 doc_id 를 들고 있는 소비자가 헛번호가 된다. - patch 판정 문단이 규칙의 patch 조항을 인용하지 않아 논거가 유리해 보였다. "문면상으로는 minor 이고 그럼에도 X 를 이유로 patch" 로 고쳤다. - 머리글 수치(11,867 / 65,764 / 11.8%)가 이 저장소에서 재현 불가라 출처를 이슈 #239 로 밝히고, 본문 실측(15.0%)과 기준이 다름을 명시했다. - SQL 을 어디에 대고 돌리는지 실행 줄을 붙였다. 검증: 워크스페이스 1301 passed / 0 failed / 64 ignored, clippy --workspace --all-targets -- -D warnings 무경고. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
217 lines
12 KiB
TOML
217 lines
12 KiB
TOML
[workspace]
|
|
resolver = "3"
|
|
members = [
|
|
"crates/kebab-core",
|
|
"crates/kebab-config",
|
|
"crates/kebab-source-fs",
|
|
"crates/kebab-parse-md",
|
|
"crates/kebab-chunk",
|
|
"crates/kebab-store-sqlite",
|
|
"crates/kebab-store-vector",
|
|
"crates/kebab-search",
|
|
"crates/kebab-embed-local",
|
|
"crates/kebab-embed-ollama",
|
|
"crates/kebab-llm-local",
|
|
"crates/kebab-rag",
|
|
"crates/kebab-app",
|
|
"crates/kebab-cli",
|
|
"crates/kebab-eval",
|
|
"crates/kebab-parse-image",
|
|
"crates/kebab-parse-pdf",
|
|
"crates/kebab-mcp",
|
|
"crates/kebab-parse-code",
|
|
"crates/kebab-nli",
|
|
]
|
|
|
|
[workspace.package]
|
|
edition = "2024"
|
|
rust-version = "1.85"
|
|
license = "MIT OR Apache-2.0"
|
|
repository = "https://github.com/altair823/kebab"
|
|
version = "0.33.1" # v0.33.1 — #239 얇은 검출 박스 하나가 이미지 OCR 전체를 날리던 문제. paddle-onnx rec 세션은 입력 폭이 4 이하면 특징맵이 0 열로 접혀 실패하는데(타임스텝 = ceil((w-4)/8)), 그 오류가 recognize 의 ? 를 타고 나가 같은 이미지에서 이미 인식해 둔 박스까지 전부 버렸다. REC_MIN_WIDTH = 5 미만은 세션에 넣지 않고 빈 문자열로 돌려보낸다. 색인은 성공으로 끝나므로 검색이 0 건 나올 때까지 안 드러나던 조용한 손실. 실측(도그푸딩 이미지 240 개 파일, 같은 모델·설정): 실패 36/240 → 0/240, 되살아난 본문 11,747 자, 원래 성공하던 204 장은 글자 수 변화 0. 곁다리로 PDF OCR 실패 노트를 {e:#} 로 바꿔 잘려 나가던 ORT 원인을 보존. patch 로 판정한 근거 — 사용자 표면(서브커맨드/플래그/config 키/wire)이 하나도 안 바뀌었다. parser_version cascade(image-meta-v1 → v2, pdf-text-v2 → v3)가 있어 minor 로 볼 여지가 있으나, v0.33.0 의 minor 트리거는 cascade 단독이 아니라 신규 config 키와 짝이었다. 다만 재색인 비용은 minor 급으로 드니(doc_id 전면 교체 + store 재작성) release notes 의 upgrade 절차 참조. 도그푸딩: HOTFIXES 2026-08-28. 상세 docs/release-notes/v0.33.1-draft.md — CLAUDE.md §Release
|
|
|
|
# pre-v0.18 workspace-wide cleanup: enable clippy::pedantic group with
|
|
# intentional allow-list. The allowed lints are either cosmetic (doc style),
|
|
# informational (function size), or carry intentional truncation we accept
|
|
# (numeric casts in tokenizer/ONNX inputs, hash modular reduction, etc).
|
|
[workspace.lints.clippy]
|
|
pedantic = { level = "warn", priority = -1 }
|
|
# Intentional u32 ↔ i64 casts in kebab-nli (ONNX i64 inputs from tokenizer u32 ids).
|
|
# u64 ↔ usize across kebab-store-sqlite row counts. Wide truncation is auditable
|
|
# at use site, not lint-wide.
|
|
cast_possible_truncation = "allow"
|
|
cast_possible_wrap = "allow"
|
|
cast_sign_loss = "allow"
|
|
cast_precision_loss = "allow"
|
|
# Doc markdown style is cosmetic; we run rustdoc on demand.
|
|
doc_markdown = "allow"
|
|
missing_errors_doc = "allow"
|
|
missing_panics_doc = "allow"
|
|
# Informational only — splitting a long pipeline function isn't always cleaner.
|
|
too_many_lines = "allow"
|
|
# `Foo::default()` is concise and idiomatic here; `<Foo as Default>::default()`
|
|
# adds noise without surfacing intent.
|
|
default_trait_access = "allow"
|
|
# Module name prefix on public items keeps the wire/log surface readable
|
|
# (`refusal_reason::no_chunks` etc).
|
|
module_name_repetitions = "allow"
|
|
# We use `#[must_use]` deliberately on public results, not blanket.
|
|
must_use_candidate = "allow"
|
|
# `String` arg sometimes signals "I'll consume this" — let signature decide.
|
|
needless_pass_by_value = "allow"
|
|
# Idiomatic single-line bindings stay; let-else expansion isn't always clearer.
|
|
manual_let_else = "allow"
|
|
# `use` after `let` is a common kebab pattern (scoped imports next to use site).
|
|
items_after_statements = "allow"
|
|
# Naming pairs like `chunk_id` / `chunks_id` are intentional domain terms.
|
|
similar_names = "allow"
|
|
# `iter.map(format!).collect::<String>()` is idiomatic when the per-element
|
|
# string is genuinely independent — `fold` only wins on accumulation patterns.
|
|
format_collect = "allow"
|
|
# Exhaustive `match` with explicit variant arms (vs `_`) catches future
|
|
# variant additions at compile time (kebab core's `RefusalReason` pattern).
|
|
match_wildcard_for_single_variants = "allow"
|
|
# Copy types under `&self` keep call-site discipline; auto-deref noise > tiny perf gain.
|
|
trivially_copy_pass_by_ref = "allow"
|
|
# `unnecessary_wraps` flags helpers that could drop `Result`, but keeping the
|
|
# Result allows future error variants without churning callers.
|
|
unnecessary_wraps = "allow"
|
|
# NLI score / RRF fusion / similarity threshold comparisons are intentional —
|
|
# floats live in the `[0, 1]` band and are compared with explicit thresholds.
|
|
float_cmp = "allow"
|
|
# File-extension dispatch is keyed on ASCII conventions; case sensitivity
|
|
# is part of the spec for `.md`, `.pdf`, etc.
|
|
case_sensitive_file_extension_comparisons = "allow"
|
|
# Config / opts structs intentionally bundle boolean flags (ingest options,
|
|
# search modes, etc) — splitting them into enums would obscure the wire shape.
|
|
struct_excessive_bools = "allow"
|
|
# `bytecount` crate would be a new dep just for one-off ASCII counts.
|
|
naive_bytecount = "allow"
|
|
# `#[ignore]` annotations on tests document via the test name + nearby comment.
|
|
ignore_without_reason = "allow"
|
|
# `format!` push patterns are common in CLI output paths;
|
|
# `write!` rewrite needs a verified-equal benchmark before swapping.
|
|
format_push_string = "allow"
|
|
# Builder-style `with_*` methods return `Self`; the existing `#[must_use]`
|
|
# discipline lives on aggregate constructors, not every chainable setter.
|
|
return_self_not_must_use = "allow"
|
|
# Match arms grouped by side-effect over body equality (e.g. snake_case wire
|
|
# label tables) — fanning them out keeps adding a new variant trivial.
|
|
match_same_arms = "allow"
|
|
# Remaining style-only warnings: trailing `continue` is sometimes clearer than
|
|
# rewriting, `_x` underscored bindings document intent at the use site, and
|
|
# `!(a == b)` reads better than `a != b` when paired with a complementary check.
|
|
needless_continue = "allow"
|
|
used_underscore_binding = "allow"
|
|
nonminimal_bool = "allow"
|
|
# Other one-off cosmetic items: large literal formatting, doc link quoting,
|
|
# `Clone::clone_from` swap, `str::replace` chaining, `Iterator::any` ergonomics.
|
|
unreadable_literal = "allow"
|
|
many_single_char_names = "allow"
|
|
doc_link_with_quotes = "allow"
|
|
assigning_clones = "allow"
|
|
collapsible_str_replace = "allow"
|
|
trivial_regex = "allow"
|
|
elidable_lifetime_names = "allow"
|
|
range_plus_one = "allow"
|
|
explicit_iter_loop = "allow"
|
|
implicit_hasher = "allow"
|
|
ref_option = "allow"
|
|
|
|
[workspace.dependencies]
|
|
anyhow = "1"
|
|
thiserror = "2"
|
|
serde = { version = "1", features = ["derive"] }
|
|
serde_json = "1"
|
|
serde_yaml_ng = "0.10"
|
|
time = { version = "0.3", features = ["serde", "macros", "formatting", "parsing"] }
|
|
uuid = { version = "1", features = ["v7", "serde"] }
|
|
blake3 = "1"
|
|
tracing = "0.1"
|
|
# `bundled` ships SQLite source so the workspace doesn't depend on a
|
|
# system libsqlite3 (matches the kebab-store-sqlite feature set).
|
|
rusqlite = { version = "0.32", features = ["bundled"] }
|
|
globset = "0.4"
|
|
tempfile = "3"
|
|
proptest = "1"
|
|
lopdf = "0.32"
|
|
# fastembed-rs ships ONNX runtime via the `ort-download-binaries` feature
|
|
# in its default set (which also pulls `hf-hub` for first-run model
|
|
# downloads). Pinned to the 4.x line per task p3-2 (current 5.x release
|
|
# remains untested for this workspace).
|
|
fastembed = "4.9"
|
|
# LanceDB embedded vector store (P3-3). 0.23.x pulls arrow / arrow-array /
|
|
# arrow-schema 56.x transitively (via lance 1.0); the kebab-store-vector
|
|
# crate matches that major to share the same Arrow types without a
|
|
# re-export adapter.
|
|
lancedb = { version = "0.23", default-features = false }
|
|
arrow = "56"
|
|
arrow-array = "56"
|
|
arrow-schema = "56"
|
|
tokio = { version = "1", features = ["rt", "macros"] }
|
|
futures = "0.3"
|
|
# Strict citation-marker extraction in kebab-rag (P4-3) needs a single regex
|
|
# pass; pulled into the workspace deps so future crates can share the
|
|
# same major.
|
|
regex = "1"
|
|
# MCP (Model Context Protocol) SDK. server + macros + transport-io provide
|
|
# stdio JSON-RPC transport for `kebab-mcp` (p9-fb-30). schemars feature
|
|
# exposes the derive macro used by tool input schemas.
|
|
rmcp = { version = "1.6", default-features = false, features = ["server", "macros", "transport-io", "schemars"] }
|
|
# Dev-only HTTP mock server for kebab-llm-local Ollama adapter tests. Requires
|
|
# a tokio runtime to host its mock server (the runtime adapter crate stays
|
|
# sync via reqwest::blocking — wiremock is dev-only there).
|
|
wiremock = "0.6"
|
|
base64 = "0.22"
|
|
# Pure-Rust git library for repo metadata detection (kebab-parse-code).
|
|
# No `git` binary required. Default features include thread-safety + most
|
|
# object-reading capabilities needed for HEAD name + commit SHA queries.
|
|
gix = { version = "0.70", default-features = false, features = ["revision"] }
|
|
# Rust source parsing for code ingest (kebab-parse-code, p10-1A-2). The
|
|
# chunker stays tree-sitter-free — AST work is parser-side per design §6.3.
|
|
tree-sitter = "0.26"
|
|
tree-sitter-rust = "0.24"
|
|
# Python / TS / JS grammars for code ingest (kebab-parse-code, p10-1B).
|
|
tree-sitter-python = "0.25.0"
|
|
tree-sitter-typescript = "0.23.2"
|
|
tree-sitter-javascript = "0.25.0"
|
|
# Go grammar for code ingest (kebab-parse-code, p10-1C-Go).
|
|
tree-sitter-go = "0.25.0"
|
|
# JVM family grammars for code ingest (kebab-parse-code, p10-1C-JK).
|
|
tree-sitter-java = "0.23.5"
|
|
tree-sitter-kotlin-ng = "1.1.0" # bare tree-sitter-kotlin requires ts <0.23; -ng uses tree-sitter-language 0.1 (ts 0.26 compat)
|
|
# C/C++ family grammars for code ingest (kebab-parse-code, p10-1D).
|
|
tree-sitter-c = "0.24.2"
|
|
tree-sitter-cpp = "0.23.4"
|
|
# fb-41 PR-9 (kebab-nli): mDeBERTa-v3 XNLI verifier deps. Versions match
|
|
# the fastembed 4.9 transitive set so the ONNX Runtime + tokenizer stack
|
|
# stays single-versioned across the workspace. ort `default-features=false`
|
|
# drops the bundled binary downloader (fastembed already provides one);
|
|
# tokenizers `default-features=false, onig` swaps the default `esaxx` regex
|
|
# backend for `onig` so the build doesn't need libstdc++ headers (verified
|
|
# via PR-9a pre-flight: SentencePiece tokenizer.json loads + KR/EN encode).
|
|
# hf-hub uses `ureq + rustls-tls` to stay aligned with kebab-embed-local's
|
|
# pure-Rust TLS stack.
|
|
ort = { version = "=2.0.0-rc.9", default-features = false, features = ["ndarray"] }
|
|
tokenizers = { version = "0.21", default-features = false, features = ["onig"] }
|
|
hf-hub = { version = "0.4", default-features = false, features = ["ureq", "rustls-tls"] }
|
|
ndarray = "0.16"
|
|
# Korean morphological tokenizer (FTS v0.20.x, §6.1). lindera-ko-dic bundles
|
|
# the KO-DIC dictionary as an embedded blob via the embed-ko-dic feature.
|
|
lindera = "3"
|
|
lindera-ko-dic = "3"
|
|
|
|
# Disk-footprint trim for dev / test builds. Codegen, opt-level, and
|
|
# behavior are unchanged — only DWARF debug info is reduced (line
|
|
# numbers kept, column numbers dropped) and split into separate
|
|
# `.dwo` files. backtrace stays useful (function + line). release
|
|
# profile is untouched, so CI / `--release` runs are byte-identical
|
|
# to upstream defaults.
|
|
[profile.dev]
|
|
debug = "line-tables-only"
|
|
split-debuginfo = "unpacked"
|
|
|
|
[profile.test]
|
|
debug = "line-tables-only"
|
|
split-debuginfo = "unpacked"
|