review(p7-1): 회차 1 지적 반영
- Cargo.toml: 사용하지 않는 deps 제거 (`kebab-config`, `thiserror`, `pdf-extract`, dev `tempfile` / `serde_json` / `serde`). 특히 `pdf-extract` 가 끌어오던 transitive ~150 crate (pom, postscript, type1-encoding-parser, adobe-cmap-parser, euclid, chrono, md5, linked-hash-map …) 가 모두 사라짐. lopdf 만 남음. - info.rs: BOM 없는 PDFDocEncoded Title 디코드 버그 수정. `from_utf8_lossy` 는 0x80–0xFF 를 U+FFFD 로 치환해 "Café" 같은 레거시 타이틀을 망가뜨림. byte → `char` 직접 캐스팅 (Latin-1 디코더) 로 교체. 회귀 테스트 `info_dict_title_pdfdocencoding_latin1_high_bytes_decoded` 추가. - info.rs: 모듈 doc 의 "Latin-1 superset" 부정확 표현 정정 — PDFDocEncoding 은 0x18–0x1F / 0x80–0x9F 영역에서 Latin-1 과 다름. - lib.rs: `saturating_sub(1)` 가 page=0 케이스를 silent 흡수하던 부분에 `debug_assert!` 추가. release 는 saturating fallback 유지 (panic 보다 garbled order 가 운영에 유리). - tests: UTF-16 surrogate pair 커버리지 갭 보완 — 🥙 (U+1F959) 가 포함된 타이틀로 `String::from_utf16_lossy` 의 페어-결합 경로 검증. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -177,6 +177,43 @@ fn info_dict_title_utf16be_bom_decoded() {
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn info_dict_title_utf16be_surrogate_pair_decoded() {
|
||||
// 🥙 (U+1F959 STUFFED FLATBREAD) sits in the supplementary plane,
|
||||
// so encoding it as UTF-16BE produces a surrogate pair (D83E DD59).
|
||||
// BMP-only inputs would never exercise the pair-joining path of
|
||||
// `String::from_utf16_lossy` — this asserts that path round-trips.
|
||||
let info = InfoDict {
|
||||
title: Some(utf16be_bom("케밥 🥙 문서")),
|
||||
producer: None,
|
||||
creator: None,
|
||||
};
|
||||
let bytes = build_text_pdf_with_info(&[Some("body")], &info);
|
||||
let fx = fixture_for("docs/emoji-title.pdf", &bytes);
|
||||
let doc = PdfTextExtractor::new()
|
||||
.extract(&fx.ctx(), &bytes)
|
||||
.expect("PDF with surrogate-pair Title must extract");
|
||||
assert_eq!(doc.title, "케밥 🥙 문서");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn info_dict_title_pdfdocencoding_latin1_high_bytes_decoded() {
|
||||
// BOM-less PDFDocEncoded title with a high-byte char (0xE9 = 'é').
|
||||
// `from_utf8_lossy` would have replaced this with U+FFFD; the
|
||||
// byte-as-char path keeps it intact.
|
||||
let info = InfoDict {
|
||||
title: Some(b"Caf\xE9".to_vec()),
|
||||
producer: None,
|
||||
creator: None,
|
||||
};
|
||||
let bytes = build_text_pdf_with_info(&[Some("body")], &info);
|
||||
let fx = fixture_for("docs/cafe-title.pdf", &bytes);
|
||||
let doc = PdfTextExtractor::new()
|
||||
.extract(&fx.ctx(), &bytes)
|
||||
.expect("PDF with Latin-1 Title must extract");
|
||||
assert_eq!(doc.title, "Café");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn info_dict_title_falls_back_to_filename_when_missing() {
|
||||
let bytes = build_text_pdf(&[Some("body")]);
|
||||
|
||||
Reference in New Issue
Block a user