Instructions to use buddhist-nlp/qwen35-mitra-dictionary-annotation with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use buddhist-nlp/qwen35-mitra-dictionary-annotation with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="buddhist-nlp/qwen35-mitra-dictionary-annotation") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("buddhist-nlp/qwen35-mitra-dictionary-annotation") model = AutoModelForMultimodalLM.from_pretrained("buddhist-nlp/qwen35-mitra-dictionary-annotation", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use buddhist-nlp/qwen35-mitra-dictionary-annotation with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "buddhist-nlp/qwen35-mitra-dictionary-annotation" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "buddhist-nlp/qwen35-mitra-dictionary-annotation", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/buddhist-nlp/qwen35-mitra-dictionary-annotation
- SGLang
How to use buddhist-nlp/qwen35-mitra-dictionary-annotation with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "buddhist-nlp/qwen35-mitra-dictionary-annotation" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "buddhist-nlp/qwen35-mitra-dictionary-annotation", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "buddhist-nlp/qwen35-mitra-dictionary-annotation" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "buddhist-nlp/qwen35-mitra-dictionary-annotation", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use buddhist-nlp/qwen35-mitra-dictionary-annotation with Docker Model Runner:
docker model run hf.co/buddhist-nlp/qwen35-mitra-dictionary-annotation
qwen35-mitra-dictionary-annotation
The word aligner behind MITRA-dict, a family of attested Sanskrit to English, Tibetan and Chinese dictionaries.
Given a Sanskrit sentence, segmented and lemmatized with ByT5-Sanskrit, and its translation, the model copies the translation and wraps the span that renders each Sanskrit word in that word's numbered tag. Words without a clear equivalent are left untagged.
Prompt
Insert the numbered tags into the target line, marking the word(s) that translate each numbered source word. Copy the target verbatim and omit numbers without a clear equivalent.
SANSKRIT: <1>evam</1> <2>ca</2> <3>tat</3> <4>karaṇīyam</4>
LEMMAS: <1>evam</1> <2>ca</2> <3>tad</3> <4>kṛ</4>
TIBETAN: དེ་ནི་དེ་ལྟ་བུར་བྱ་སྟེ།
OUTPUT:
Output:
དེ་ནི་<1>དེ་ལྟ་བུར</1>་<4>བྱ</4>་སྟེ།
The target line is TIBETAN:, ENGLISH: or CHINESE: (followed by a CHINESE_SEGMENTED: line with the word-segmented Chinese). Tibetan is given in Unicode. Use greedy decoding and stop at the first newline.
Training
Full fine-tuning of MITRA Qwen3.5-9B (continued pretraining on Sanskrit, Tibetan, Buddhist Chinese and Pāli) for three epochs on 9,366 sentence pairs labelled by a commercial LLM. Directions are Sanskrit to Tibetan, Chinese and English, plus Chinese to Tibetan and English to Tibetan.
Evaluation
Link-level F1 against 1,489 held-out sentence pairs labelled by the same teacher: Sanskrit to Tibetan 88.9, Sanskrit to Chinese 89.2, Sanskrit to English 86.9. These figures measure agreement with the teacher, not correctness.
Format
The weights use the Qwen3_5ForConditionalGeneration layout, which vLLM requires. The vision tower is unused.
Links
Data and evaluation protocol: https://github.com/dharmamitra/dharmamitra-lexicon
Interactive dictionary: https://lexicon.dharmamitra.org
- Downloads last month
- 177
Model tree for buddhist-nlp/qwen35-mitra-dictionary-annotation
Base model
buddhist-nlp/mitra-qwen35-base-stage2