혼합 출처 KB(위키+jira 등)에서 색인은 전부 하되 질의 시 출처로 좁히는 provenance 레버. 전역 trust 곱셈가중(weighted-RRF)은 A/B 에서 반증(θ=0.85 만으로 incident MRR 0.918→0.340 절벽, 점수 압축) — 필터가 see-saw 없는 올바른 레버. - config [[workspace.sources]] (각 id/root/exclude/trust_level/source_type); 단일 root 는 implicit `default` source 로 정규화. validate: id 유일·비어있지 않음. - config schema v3→v4 (step_3_to_4, root→[[workspace.sources]] id=default 미러, 멱등) - V014 documents.source_id 컬럼+인덱스 (additive, DEFAULT 'default', 재색인 0) - Metadata.source_id + BodyHints trust precedence(frontmatter > source 기본값 > Primary) - ingest: --root 미지정 시 resolved_sources() 순회 + doc 마다 source_id/trust stamp - 검색 SearchFilters.source_type/source_id → lexical + vector 두 site (IN, OR) - CLI kebab search --source <id> / --source-type <type> (repeatable/comma-sep) 도그푸딩(620 doc, jira400+wiki220): --source wiki 로 개념 질의 MRR 0.780→0.810, --source jira 로 incident 0.918→0.975. trust precedence 실측(jira=secondary 기본값). version bump 0.28.0 → 0.29.0 (신규 CLI flag + config 키 + V014 migration → minor). follow-up: MCP search 필터 미노출 · kebab list source_id 미표시 · RAG provenance 라벨. 자세한 내용: tasks/HOTFIXES.md (2026-06-21), docs/release-notes/v0.29.0-draft.md. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012Mc6W1fgsrbFKTsqA6P8La
126 lines
4.2 KiB
Rust
126 lines
4.2 KiB
Rust
//! Integration: tools/call name=ingest_file → ingest_report.v1.
|
|
|
|
use std::fs;
|
|
|
|
use kebab_config::Config;
|
|
use kebab_mcp::{KebabAppState, KebabHandler};
|
|
use rmcp::model::RawContent;
|
|
|
|
#[tokio::test]
|
|
async fn ingest_file_tool_returns_ingest_report_v1() {
|
|
let dir = tempfile::tempdir().unwrap();
|
|
let workspace = dir.path().join("notes");
|
|
let data = dir.path().join("data");
|
|
fs::create_dir_all(&workspace).unwrap();
|
|
fs::create_dir_all(&data).unwrap();
|
|
|
|
let mut cfg = Config::defaults();
|
|
cfg.workspace.root = Some(workspace.to_string_lossy().into_owned());
|
|
cfg.storage.data_dir = data.to_string_lossy().into_owned();
|
|
cfg.models.embedding.provider = "none".to_string();
|
|
cfg.models.embedding.dimensions = 0;
|
|
|
|
let src = dir.path().join("doc.md");
|
|
fs::write(&src, "# Title\n\nbody.").unwrap();
|
|
|
|
let state = KebabAppState::new(cfg, None);
|
|
let handler = KebabHandler::new(state);
|
|
|
|
let result = tokio::task::spawn_blocking({
|
|
let state = handler.state().clone();
|
|
let path = src.to_string_lossy().into_owned();
|
|
move || {
|
|
kebab_mcp::tools::ingest_file::handle(
|
|
&state,
|
|
kebab_mcp::tools::ingest_file::IngestFileInput { path },
|
|
)
|
|
}
|
|
})
|
|
.await
|
|
.unwrap();
|
|
|
|
assert!(!result.is_error.unwrap_or(false), "{result:?}");
|
|
let text = match &result.content.first().unwrap().raw {
|
|
RawContent::Text(t) => &t.text,
|
|
other => panic!("expected text content, got {other:?}"),
|
|
};
|
|
let v: serde_json::Value = serde_json::from_str(text).unwrap();
|
|
assert_eq!(
|
|
v.get("schema_version").and_then(|s| s.as_str()),
|
|
Some("ingest_report.v1")
|
|
);
|
|
assert_eq!(v.get("new").and_then(serde_json::Value::as_u64), Some(1));
|
|
}
|
|
|
|
#[tokio::test]
|
|
async fn ingest_file_tool_idempotent_on_second_call() {
|
|
let dir = tempfile::tempdir().unwrap();
|
|
let workspace = dir.path().join("notes");
|
|
let data = dir.path().join("data");
|
|
std::fs::create_dir_all(&workspace).unwrap();
|
|
std::fs::create_dir_all(&data).unwrap();
|
|
|
|
let mut cfg = kebab_config::Config::defaults();
|
|
cfg.workspace.root = Some(workspace.to_string_lossy().into_owned());
|
|
cfg.storage.data_dir = data.to_string_lossy().into_owned();
|
|
cfg.models.embedding.provider = "none".to_string();
|
|
cfg.models.embedding.dimensions = 0;
|
|
|
|
let src = dir.path().join("doc.md");
|
|
std::fs::write(&src, "# A\n\nbody.").unwrap();
|
|
|
|
let state = kebab_mcp::KebabAppState::new(cfg, None);
|
|
let handler = kebab_mcp::KebabHandler::new(state);
|
|
|
|
// First call.
|
|
let r1 = tokio::task::spawn_blocking({
|
|
let state = handler.state().clone();
|
|
let path = src.to_string_lossy().into_owned();
|
|
move || {
|
|
kebab_mcp::tools::ingest_file::handle(
|
|
&state,
|
|
kebab_mcp::tools::ingest_file::IngestFileInput { path },
|
|
)
|
|
}
|
|
})
|
|
.await
|
|
.unwrap();
|
|
assert!(!r1.is_error.unwrap_or(false));
|
|
let text1 = match &r1.content.first().unwrap().raw {
|
|
rmcp::model::RawContent::Text(t) => &t.text,
|
|
other => panic!("expected text, got {other:?}"),
|
|
};
|
|
let v1: serde_json::Value = serde_json::from_str(text1).unwrap();
|
|
assert_eq!(v1.get("new").and_then(serde_json::Value::as_u64), Some(1));
|
|
|
|
// Second call — same content, expect unchanged=1.
|
|
let r2 = tokio::task::spawn_blocking({
|
|
let state = handler.state().clone();
|
|
let path = src.to_string_lossy().into_owned();
|
|
move || {
|
|
kebab_mcp::tools::ingest_file::handle(
|
|
&state,
|
|
kebab_mcp::tools::ingest_file::IngestFileInput { path },
|
|
)
|
|
}
|
|
})
|
|
.await
|
|
.unwrap();
|
|
assert!(!r2.is_error.unwrap_or(false));
|
|
let text2 = match &r2.content.first().unwrap().raw {
|
|
rmcp::model::RawContent::Text(t) => &t.text,
|
|
other => panic!("expected text, got {other:?}"),
|
|
};
|
|
let v2: serde_json::Value = serde_json::from_str(text2).unwrap();
|
|
assert_eq!(
|
|
v2.get("new").and_then(serde_json::Value::as_u64),
|
|
Some(0),
|
|
"{v2:?}"
|
|
);
|
|
assert_eq!(
|
|
v2.get("unchanged").and_then(serde_json::Value::as_u64),
|
|
Some(1),
|
|
"{v2:?}"
|
|
);
|
|
}
|