guard-ft-v19 (vocabulary-trimmed)
GUARD-FT round 19 (AgentDocs' EN+FR legal-document PII token classifier). This repo's main now
carries the VOCABULARY-TRIMMED export: XLM-R-large's 250,002-piece vocabulary reduced to the 53,751
pieces the EN+FR training corpus actually used (plus every byte-level/single-character fallback
piece), tools/guard-train/trim_vocab.py. Same checkpoint, same weights on every kept row — this
is row-selection on the embedding matrix, not a retrain, and measured byte-identical to the
untrimmed export's masking decisions across a 259-document evaluation corpus.
- Weights: ~2.1 GB (untrimmed) -> ~1.4 GB (this revision), now a single self-contained
model.onnx(under the 2 GB protobuf limit, so no companionmodel.onnx_dataany more). - The untrimmed original remains reachable at the pinned historical revision
09451885cc53531ecd28aa3bbc1be7b7986ffb9eif you need the full 250,002-piece vocabulary. - Sibling exports:
guard-ft-v19-trimmed-fp16(CUDA-only, ~716 MB) andguard-ft-v19-trimmed-int8(CPU-only, ~360 MB, measurable precision trade-off). - Labels: PERSON, ORG, PUBLIC_BODY, ROLE, ADDRESS, PLACE, DOB, CONTACT (BIO scheme)
- Downloads last month
- 15
Model tree for gauthierrobert2/guard-ft-v19
Base model
FacebookAI/xlm-roberta-large