Instructions to use datumo/E-star-12B-base with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use datumo/E-star-12B-base with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="datumo/E-star-12B-base") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("datumo/E-star-12B-base") model = AutoModelForCausalLM.from_pretrained("datumo/E-star-12B-base", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use datumo/E-star-12B-base with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "datumo/E-star-12B-base" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "datumo/E-star-12B-base", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/datumo/E-star-12B-base
- SGLang
How to use datumo/E-star-12B-base with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "datumo/E-star-12B-base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "datumo/E-star-12B-base", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "datumo/E-star-12B-base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "datumo/E-star-12B-base", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use datumo/E-star-12B-base with Docker Model Runner:
docker model run hf.co/datumo/E-star-12B-base
datumo/E-Star-12B-v2-Base
- ์์ ์: Selectstar Eval Team
- ์์ฑ์ผ: 2026-05-22
- ์ํ: Active
- ๋ฌธ์์ ๋ฒ์ : v0.1
- Hugging Face ID:
datumo/E-Star-12B-v2-Base
๋น์์ ์ ์ฌ์ฉ ์ ์ฉ(Non-Commercial Use Only)
๋ณธ ๋ชจ๋ธ์ Gemma 3์ ํ์ ๋ชจ๋ธ์ด๋ฏ๋ก Gemma Terms of Use๊ฐ ์ ์ฉ๋ฉ๋๋ค. Selectstar๊ฐ ๋ณด์ ํ ์์ ๋ถ์๋ CC BY-NC 4.0 ๊ธฐ๋ฐ์ ์ถ๊ฐ ๋น์์ ์กฐ๊ฑด์ด ์ ์ฉ๋ฉ๋๋ค. ๋ ์กฐ๊ฑด์ ๋ชจ๋ ์ค์ํด์ผ ํ๋ฉฐ, ์์ ์ ์ด์ฉ์ Selectstar์ ๋ณ๋ ๊ณ์ฝ์ด ํ์ํฉ๋๋ค.
1. ๋ชจ๋ธ ์ค๋ช
1.1 ๋ชจ๋ธ ์ํคํ ์ฒ / ํ๋ผ๋ฏธํฐ ํฌ๊ธฐ
| ํญ๋ชฉ | ๋ด์ฉ |
|---|---|
| ๋ฒ ์ด์ค ๋ชจ๋ธ | google/gemma-3-12b-it |
| ๋ชจ๋ธ ๊ณ์ด | Gemma 3 |
| ์ํคํ ์ฒ | Transformer ๊ธฐ๋ฐ Decoder ๊ณ์ด |
| ํ๋ผ๋ฏธํฐ ํฌ๊ธฐ | ์ฝ 12B |
| ํ์ต ๋ฐฉ์ | Full Fine-Tuning ๊ธฐ๋ฐ SFT |
| ๋ชจ๋ธ ์ ํ | ํ๊ตญ์ด ๋ฃจ๋ธ๋ฆญ ๊ธฐ๋ฐ Evaluation Model |
| ์ฃผ์ ์ถ๋ ฅ | feedback โ highlight โ decision |
| ์ฃผ์ ์ธ์ด | ํ๊ตญ์ด |
๋ณธ ๋ชจ๋ธ์ ์ฃผ์ด์ง ๋ฌธ์ , ๋ชจ๋ธ ์๋ต, ํ๊ฐ ๊ธฐ์ค๊ณผ ์ ์๋ณ ๋ฃจ๋ธ๋ฆญ์ ํด์ํ ๋ค ํ๊ฐ ์ค๋ช , ํต์ฌ ๊ทผ๊ฑฐ ๊ตฌ๊ฐ ๋ฐ ์ต์ข ์ ์๋ฅผ ๊ตฌ์กฐํ๋ ํ์์ผ๋ก ์์ฑํ๋ค.
Gemma 3 12B๋ ๋ฉํฐ๋ชจ๋ฌ ์ ๋ ฅ์ ์ง์ํ๋ ๋ชจ๋ธ ๊ณ์ด์ด์ง๋ง, ๋ณธ ๋ชจ๋ธ์ ํ์ต ๋ฐ ํ๊ฐ ๋ฒ์๋ ํ์ฌ ๋ฌธ์ ๊ธฐ์ค ํ ์คํธ ์ ๋ ฅ์ผ๋ก ์ ํ๋๋ค. ์ด๋ฏธ์ง ์ ๋ ฅ ์ฑ๋ฅ์ ๊ฒ์ฆ๋์ง ์์๋ค.
1.2 ๋ชจ๋ธ ์ ์์ ์ฌ์ฉํ GitHub ์ ์ฅ์
| ํญ๋ชฉ | ๋ด์ฉ |
|---|---|
| ํ์ต ์ ์ฅ์ | ํ์ธ ํ์ |
| ํ๊ฐ ์ ์ฅ์ | ํ์ธ ํ์ |
| ์ถ๋ก ยท์๋น ์ ์ฅ์ | ํ์ธ ํ์ |
| ํ์ต commit ๋๋ tag | ํ์ธ ํ์ |
| Config ๊ฒฝ๋ก | ํ์ธ ํ์ |
์ค์ ๋ชจ๋ธ ์ ์์ ์ฌ์ฉํ ์ ์ฅ์ URL๊ณผ ์ฌํ ๊ฐ๋ฅํ commit hash ๋๋ release tag๋ฅผ ์ถ๊ฐํด์ผ ํ๋ค. ๋ด๋ถ ์ ์ฅ์๋ผ๋ฉด ์ ๊ทผ ๊ถํ๊ณผ ๋ด๋น ์กฐ์ง์ ํจ๊ป ๊ธฐ์ฌํ๋ค.
1.3 ๋ชจ๋ธ ๋ฒ์ ์ ๋ณด
| ํญ๋ชฉ | ๋ด์ฉ |
|---|---|
| Hugging Face ์ ์ฅ์๋ช | E-Star-12B-v2-Base |
| ๋ฌธ์์ ๋ฒ์ | v0.1 |
| ๋ฒ์ ์ค๋ช | K2-Feedback ๊ธฐ๋ฐ 3๋จ๊ณ ํํฐ๋ง ๋ฐ์ดํฐ 6,311๊ฐ๋ก ํ์ตํ ์ด๊ธฐ Base ๋ฒ์ |
| ํ์ต ์ฒดํฌํฌ์ธํธ | ํ์ธ ํ์ |
์ ์ฅ์๋ช
์ v2์ ๋ฌธ์ ๋ฒ์ v0.1์ ์๋ฏธ๊ฐ ํผ์ฌ๋์ด ์๋ค. v2๊ฐ ์ ํ ์ธ๋์ด๊ณ v0.1์ด ์ฒดํฌํฌ์ธํธ ๋ฒ์ ์ด๋ผ๋ฉด ์ด๋ฅผ ๋ช
์ํ๊ณ , ๊ทธ๋ ์ง ์๋ค๋ฉด ๋ชจ๋ธ๋ช
๊ณผ ๋ฒ์ ์ ํต์ผํด์ผ ํ๋ค.
1.4 ๋ชฉ์ / ์ฌ์ฉ ์ฌ๋ก
E-Star-12B-v2-Base๋ ํ๊ตญ์ด ํ๊ฒฝ์์ evaluator๊ฐ ์ฃผ์ด์ง ๋ฃจ๋ธ๋ฆญ์ ์ผ๋ง๋ ์ ํํ๊ณ ์ผ๊ด๋๊ฒ ์ ์ฉํ๋์ง ํ๊ฐํ๊ณ ์๋ํํ๊ธฐ ์ํ ๋ชจ๋ธ์ด๋ค.
| ํ๊ฐ ์ถ | ์ค๋ช |
|---|---|
| Faithfulness | ๋ชจ๋ธ ์๋ต์ด ์ ๊ณต๋ ๋ฌธ์์ ๊ทผ๊ฑฐํ๋์ง ํ๋จ |
| Context Relevancy | ๊ฒ์๋ ๋ฌธ์๊ฐ ์ฌ์ฉ์ ์ง์์ ๊ด๋ จ๋๋์ง ํ๋จ |
| Response Relevancy | ๋ชจ๋ธ ์๋ต์ด ์ฌ์ฉ์ ์ง์์ ์ ์ ํ ๋์ํ๋์ง ํ๋จ |
| Rubric Following | ์์์ ํ๊ฐ ๊ธฐ์ค๊ณผ ์ ์๋ณ ์ค๋ช ์ ์ผ๊ด๋๊ฒ ์ ์ฉํ๋์ง ํ๋จ |
์ฃผ์ ์ฌ์ฉ ์ฌ๋ก๋ ๋ค์๊ณผ ๊ฐ๋ค.
- ๊ธ์ตยท๋ฒ๋ฅ RAG ํ์ดํ๋ผ์ธ ํ์ง ํ๊ฐ
- ๋ชจ๋ธ ์๋ต์ ๋ํ 1~5์ ๋ฃจ๋ธ๋ฆญ ํ๊ฐ
- ํ๊ฐ ๊ทผ๊ฑฐ ๋ฐ ํต์ฌ ํ ์คํธ ๊ตฌ๊ฐ ์ถ์ถ
- LLM ์๋ต ํ์ง ๋ชจ๋ํฐ๋ง๊ณผ ํ๊ท ํ ์คํธ
- ์ฌ๋ ํ๊ฐ ์ด์ ์ 1์ฐจ ์๋ ํ๊ฐ
- Frontier judge ๋ชจ๋ธ ๋๋น ๋น์ฉ ํจ์จ์ ์ธ ๋ด๋ถ evaluator
1.5 ๋ชจ๋ธ ์์ฉ ๊ฐ๋ฅ์ฑ
- ๋๋ฉ์ธ๋ณ ๋ฃจ๋ธ๋ฆญ์ ์ถ๊ฐํ์ฌ ๊ธ์ตยท๋ฒ๋ฅ ์ธ ๋ถ์ผ๋ก ํ์ฅ
- ์ผ๋ฐ instruction following ๋ฐ reasoning ์๋ต ํ๊ฐ
- RAG ๊ฒ์๊ธฐยท์์ฑ๊ธฐยทํ๋กฌํํธ ๊ฐ ๋น๊ต ํ๊ฐ
- ๋ฐ์ดํฐ ํ์ง ๊ฒ์์ ํ์ต ์ํ ์ฐ์ ์์ ์ ๋ณ
- ์ ์์ ์์ฐ์ด ํผ๋๋ฐฑ์ด ํจ๊ป ํ์ํ ํ๊ฐ ํ์ดํ๋ผ์ธ
์๋ก์ด ๋๋ฉ์ธ์ ์ ์ฉํ ๋๋ ์ ๋ฌธ๊ฐ ๊ฒ์ ๋ฐ์ดํฐ๋ก ๋ณ๋ ์ฑ๋ฅ ๊ฒ์ฆ์ด ํ์ํ๋ค.
2. ๋ชจ๋ธ ์คํ ๋ฐฉ๋ฒ
2.1 ์คํ ํ๊ฒฝ
| ํญ๋ชฉ | ๋ด์ฉ |
|---|---|
| Python | ํ์ธ ํ์ |
| PyTorch | ํ์ธ ํ์ |
| Transformers | ํ์ธ ํ์ |
| TRL | ํ์ธ ํ์ |
| ์ ๋ฐ๋ | BF16 |
| ์คํ ์ฅ์น | CUDA GPU ๊ถ์ฅ |
2.2 ํ์ต ์ฝ๋ ์ค๋ํซ
์๋ ์ฝ๋๋ ์ ๊ณต๋ ์ค์ ์ ์์ฝํ ์์๋ค. ์ค์ ์ฌํ์๋ ํ์ต ์ ์ฅ์, ์ ์ฒ๋ฆฌ ์ฝ๋์ ์ ํํ ํจํค์ง ๋ฒ์ ์ด ํ์ํ๋ค.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from trl import SFTConfig, SFTTrainer
BASE_MODEL = "google/gemma-3-12b-it"
tokenizer = AutoTokenizer.from_pretrained(
BASE_MODEL,
trust_remote_code=True,
)
model = AutoModelForCausalLM.from_pretrained(
BASE_MODEL,
torch_dtype=torch.bfloat16,
trust_remote_code=True,
)
sft_config = SFTConfig(
output_dir="./outputs/e-star-12b-v2-base",
per_device_train_batch_size=1,
gradient_accumulation_steps=8,
learning_rate=1e-5,
num_train_epochs=5,
eval_strategy="epoch",
save_strategy="epoch",
load_best_model_at_end=True,
metric_for_best_model="eval_loss",
greater_is_better=False,
bf16=True,
logging_steps=10,
)
trainer = SFTTrainer(
model=model,
args=sft_config,
train_dataset=train_dataset,
eval_dataset=eval_dataset,
processing_class=tokenizer,
)
trainer.train()
num_train_epochs=5์ โ2 epoch์์ early stoppingโ์ ์๋ก ๋ค๋ฅธ ์ ๋ณด๋ค. Early stopping์ ์ฌ์ฉํ๋ค๋ฉด callback, patience์ ์ค์ ์ข
๋ฃ epoch๋ฅผ ํ์ต ๋ก๊ทธ ๊ธฐ์ค์ผ๋ก ๋ช
์ํด์ผ ํ๋ค.
2.3 ์ถ๋ก ์ฝ๋ ์ค๋ํซ
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
MODEL_ID = "datumo/E-Star-12B-v2-Base"
tokenizer = AutoTokenizer.from_pretrained(
MODEL_ID,
trust_remote_code=True,
)
model = AutoModelForCausalLM.from_pretrained(
MODEL_ID,
torch_dtype=torch.bfloat16,
device_map="auto",
trust_remote_code=True,
)
model.eval()
system_prompt = """You are a rubric evaluator.
Evaluate the response strictly according to the provided rubric.
Return exactly <feedback>, <highlight>, and <decision>."""
user_prompt = """You MUST write all output in the same language as the input.
# Data to Evaluate
### Problem
ํ ๊ณต์ฅ์์ ํ๋ฃจ์ 120๊ฐ์ ์ ํ์ ์์ฐํ๋ค. ๋ถ๋๋ฅ ์ด 5%์ผ ๋,
์ผ์ฃผ์ผ ๋์ ์์ฐ๋๋ ์ ์ ์ ํ์ ์๋?
### Model Response
ํ๋ฃจ ์ ์ ์ ํ์ 120 - (120 ร 0.05) = 114๊ฐ์ด๋ฉฐ,
์ผ์ฃผ์ผ ์ ์ ์ ํ์ 114 ร 7 = 798๊ฐ์ด๋ค.
### Optional Ground Truth
798๊ฐ
# Rubric
์ ๋ต์ ์ ํ์ฑ๊ณผ ํ์ด ๊ณผ์ ์ ๋
ผ๋ฆฌ์ ์ผ๊ด์ฑ์ 1์ ์์ 5์ ์ผ๋ก ํ๊ฐํ๋ค.
"""
messages = [
{"role": "system", "content": system_prompt},
{"role": "user", "content": user_prompt},
]
inputs = tokenizer.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_tensors="pt",
return_dict=True,
).to(model.device)
with torch.inference_mode():
outputs = model.generate(
**inputs,
max_new_tokens=2048,
do_sample=False,
)
new_tokens = outputs[0, inputs["input_ids"].shape[1]:]
print(tokenizer.decode(new_tokens, skip_special_tokens=True))
์ค์ ํ์ต์ ์ฌ์ฉํ ์ ์ฒด prompt, ์ถ๋ ฅ ํ์์ generation config๋ฅผ ํจ๊ป ์ ๊ณตํด์ผ ํ๊ฐ ๊ฒฐ๊ณผ๋ฅผ ์ฌํํ ์ ์๋ค.
3. ํ์ต ๋ฐ์ดํฐ์ ์ค๋ช
3.1 ๋ฐ์ดํฐ ์ถ์ฒ ๋ฐ ํน์ฑ
- ํ์ต ๋ฐ์ดํฐ์ : datumo/E-Star-Train-6K
- Dataset Card:
EVAL-EStar-SFT-v0.1
| ํญ๋ชฉ | ๋ด์ฉ |
|---|---|
| ์๋ ๋ฐ์ดํฐ | K2-Feedback (HAERAEHUB, 2024) |
| ์๋ ๋ฐ์ดํฐ ๊ท๋ชจ | ์ฝ 99,700๊ฐ |
| ์ต์ข ํ์ต ๋ฐ์ดํฐ | 6,311๊ฐ |
| ์ธ์ด | ํ๊ตญ์ด |
| ํ๋ | system, user, assistant |
| ์ถ๋ ฅ ๊ตฌ์กฐ | feedback โ highlight โ decision |
| ์์ฑ ์ ํ | Human+AI |
| ๋ ์ด๋ธ ๊ฒ์ฆ | ๋ณต์ ๋ชจ๋ธ ํฉ์ ๋ฐ frontier ๋ชจ๋ธ ๊ต์ฐจ ๊ฒ์ฆ |
3.2 ํํฐ๋ง ํ์ดํ๋ผ์ธ
| ๋จ๊ณ | ๊ท๋ชจ ๋ณํ | ์ฒ๋ฆฌ ๋ฐฉ๋ฒ |
|---|---|---|
| Stage 1 | 99.7K โ 26K | Qwen3-30B-A3B-Instruct-2507๊ณผ Qwen3-Next-80B-A3B-Instruct์ ์ด๊ธฐ ํฉ์ |
| Stage 2 | 26K โ 8K | Gemma ํ๋จ๊ณผ Qwen ํฉ์ ๊ฐ ์ผ์นยท๋ถ์ผ์น ์ํ ๊ท ํํ |
| Stage 3 | 8K โ 6,311 | GPT-5.2 ๋จ์ผ ํ๊ฐ, ์ํ frontier debate์ Qwen ํฉ์ ๊ต์ฐจ ๊ฒ์ฆ |
์ต์ข ๋ฐ์ดํฐ๋ ์ธ ํ๋จ ๊ฒฝ๋ก ์ค ์ต์ ๋ ๊ฒฝ๋ก ์ด์์ด ๋์ผ ๋ ์ด๋ธ์ ์ง์งํ ์ํ๋ก ๊ตฌ์ฑ๋๋ค.
3.3 ํ๊ฐ ๋ฒค์น๋งํฌ
- datumo/Feedback-Bench
- ์์ดยทํ๊ตญ์ด rubric following ํ๊ฐ
- 5์ ์ฒ๋, reference-free, debate ๊ธฐ๋ฐ ์ฌ๋ ์ด๋ธ๋ง
- datumo/Rag-Quality-Bench
- ๊ธ์ตยท๋ฒ๋ฅ domain adaptation ํ๊ฐ
- Context Relevancy, Faithfulness, Response Relevancy
3.4 ๋ฐ์ดํฐ ๋์ ํ์ธ
ํ์ต ๋ฐ์ดํฐ์ ํ๊ฐ ๋ฒค์น๋งํฌ๊ฐ ์ ์ฌํ debate ์ ์ฐจ๋ฅผ ์ฌ์ฉํ๋ฏ๋ก instruction, response, rubric, ์์ฒ ๋ฌธ์์ ์์ฑ ๋ชจ๋ธ ์กฐํฉ์ ์ค๋ณต ์ฌ๋ถ๋ฅผ ํ์ธํด์ผ ํ๋ค.
4. ํ์ต ์ค์
4.1 ์ฃผ์ ํ์ต ํ๋ผ๋ฏธํฐ
| ํญ๋ชฉ | ์ค์ |
|---|---|
| ํ์ต ๋ฐฉ์ | Full Fine-Tuning |
| ํ๋ ์์ํฌ | TRL SFTTrainer |
| Learning rate | 1e-5 |
| ์ต๋ epoch | 5 |
| ์ค์ ์ข ๋ฃ epoch | 2๋ก ๊ธฐ์ฌ๋์ด ์์ผ๋ ๋ก๊ทธ ํ์ธ ํ์ |
| Per-device batch size | 1 |
| Gradient accumulation | 8 |
| Precision | BF16 |
| Evaluation strategy | Epoch |
| Save strategy | Epoch |
| Best-model metric | Validation loss |
| LoRA | ์ฌ์ฉํ์ง ์์ |
| Optimizer / Scheduler | ํ์ธ ํ์ |
| Max sequence length | ํ์ธ ํ์ |
| Random seed | ํ์ธ ํ์ |
4.2 GPU ๋ฐ ํ์ต ์๊ฐ
| ํญ๋ชฉ | ๋ด์ฉ |
|---|---|
| GPU ๊ฐ์ | 4๊ฐ |
| GPU ๋ชจ๋ธ | ํ์ธ ํ์ |
| GPU๋น ๋ฉ๋ชจ๋ฆฌ | ํ์ธ ํ์ |
| ๋ถ์ฐ ํ์ต ๋ฐฉ์ | ํ์ธ ํ์ |
| ์ด ํ์ต ์๊ฐ | ์ฝ 1.2์๊ฐ |
GPU: 4๋ง์ผ๋ก๋ ์ฌํํ ์ ์์ผ๋ฏ๋ก A100, H100, H200 ๋ฑ ์ ํํ GPU ๋ชจ๋ธ๊ณผ ๋ฉ๋ชจ๋ฆฌ ์ฉ๋์ ์ถ๊ฐํด์ผ ํ๋ค.
5. ํ๊ฐ ๊ฒฐ๊ณผ
5.1 Feedback Bench โ ์์ด Rubric Following
| Type | Model | Pearson | Kendall ฯ | Spearman |
|---|---|---|---|---|
| Frontier | GPT-5.2 | 0.916 | 0.865 | 0.911 |
| Frontier | Sonnet-4.6 | 0.840 | 0.776 | 0.847 |
| Instruct SLM | Gemma-3-12B-IT | 0.810 | 0.725 | 0.794 |
| Instruct SLM | oss-20b | 0.844 | 0.762 | 0.839 |
| Evaluator LM | Prometheus-8x7B-v2.0 | 0.823 | 0.736 | 0.806 |
| Evaluator LM | GLIDER 3.8B | 0.678 | 0.595 | 0.688 |
| Ours | E-Star-12B-Base | 0.856 | 0.778 | 0.847 |
5.2 Ko Feedback Bench โ ํ๊ตญ์ด Rubric Following
| Type | Model | Pearson | Kendall ฯ | Spearman |
|---|---|---|---|---|
| Frontier | GPT-5.2 | 0.929 | 0.886 | 0.925 |
| Frontier | Sonnet-4.6 | 0.820 | 0.758 | 0.833 |
| Instruct SLM | Gemma-3-12B-IT | 0.653 | 0.593 | 0.661 |
| Instruct SLM | oss-20b | 0.778 | 0.704 | 0.779 |
| Evaluator LM | Prometheus-8x7B-v2.0 | 0.377 | 0.441 | 0.501 |
| Evaluator LM | GLIDER 3.8B | 0.523 | 0.487 | 0.563 |
| Ours | E-Star-12B-Base | 0.826 | 0.754 | 0.819 |
5.3 RAG Quality Bench โ ๊ธ์ตยท๋ฒ๋ฅ Domain Adaptation
| Model | LAW CR | LAW F | LAW RR | FIN CR | FIN F | FIN RR | Average |
|---|---|---|---|---|---|---|---|
| GPT-5.2 | 0.846 | 0.785 | 0.941 | 0.882 | 0.740 | 0.970 | 0.861 |
| Sonnet-4.6 | 0.910 | 0.786 | 0.872 | 0.932 | 0.845 | 0.925 | 0.878 |
| Gemma-3-12B-IT | 0.620 | 0.742 | 0.742 | 0.830 | 0.713 | 0.821 | 0.745 |
| oss-20b | 0.846 | 0.722 | 0.870 | 0.793 | 0.752 | 0.900 | 0.813 |
| Prometheus-8x7B-v2.0 | 0.392 | 0.477 | 0.772 | 0.386 | 0.240 | 0.806 | 0.512 |
| GLIDER 3.8B | 0.657 | 0.670 | 0.680 | 0.432 | 0.415 | 0.548 | 0.567 |
| E-Star-12B-Base | 0.853 | 0.730 | 0.816 | 0.835 | 0.720 | 0.880 | 0.806 |
CR์ Context Relevancy, F๋ Faithfulness, RR์ Response Relevancy๋ฅผ ์๋ฏธํ๋ค.
5.4 ๊ฒฐ๊ณผ ํด์
- E-Star๋ Feedback Bench์ Ko Feedback Bench์์ ๋ฒ ์ด์ค ๋ชจ๋ธ๋ณด๋ค ๋์ ์๊ด๋๋ฅผ ๊ธฐ๋กํ๋ค.
- RAG Quality ํ๊ท ์ 0.806์ผ๋ก Gemma-3-12B-IT์ 0.745๋ณด๋ค ๋์ง๋ง GPT-5.2, Sonnet-4.6๊ณผ oss-20b๋ณด๋ค๋ ๋ฎ๋ค.
- Faithfulness๋ ๋ฒ๋ฅ 0.730, ๊ธ์ต 0.720์ผ๋ก ๋ฒ ์ด์ค ๋ชจ๋ธ๋ณด๋ค ๋ฎ๊ฑฐ๋ ์ ์ฌํด ์งํ๋ณ ์ทจ์ฝ์ ๋ถ์์ด ํ์ํ๋ค.
5.5 ํ๊ฐ ์ฌํ ์ ๋ณด
ํ๊ฐ ์ฝ๋์ commit, ๋ฒค์น๋งํฌ revision, ๋ชจ๋ธ ๋ฒ์ , prompt, generation config, ์ถ๋ ฅ ํ์ฑ ๊ท์น, ํ๊ฐ์ผ๊ณผ ์ ๋ขฐ๊ตฌ๊ฐ์ ์ถ๊ฐํด์ผ ํ๋ค.
Evaluation report: ํ์ธ ํ์
Evaluation repository: ํ์ธ ํ์
Raw result artifact: ํ์ธ ํ์
6. ํ๊ณ
- Ko Feedback Bench์ ๊ธฐ๊ณ๋ฒ์ญ ๋ ธ์ด์ฆ๊ฐ ์ฑ๋ฅ ์ธก์ ์ ์ํฅ์ ์ค ์ ์๋ค.
- ํ์ต ๋ฐ์ดํฐ์ ๋ฒค์น๋งํฌ๊ฐ ์ ์ฌํ debate ์ ์ฐจ๋ก ๊ตฌ์ถ๋์ด ํน์ judge ํฉ์ ๊ธฐ์ค๊ณผ์ ์ ๋ ฌ์ ์ธก์ ํ ์ ์๋ค.
- ๋๋ฉ์ธ ์ ๋ฌธ๊ฐ์ ๋ ๋ฆฝ์ ์ธ human evaluation์ด ์ํ๋์ง ์์๋ค.
- Reference-free ์ค์ ์ผ๋ก ํ์ตยทํ๊ฐ๋์ด reference ํฌํจ ํ๊ฒฝ์ ๋ณ๋ ๊ฒ์ฆ์ด ํ์ํ๋ค.
- ๋ณต์กํ๊ฑฐ๋ ์์ถฉํ๋ ๋ฃจ๋ธ๋ฆญ์์๋ frontier ๋ชจ๋ธ๋ณด๋ค ํ๋จ ์ฑ๋ฅ์ด ๋ฎ์ ์ ์๋ค.
- ์ ๋ ฅ ๋ฌธ์ ์์ context ๊ธธ์ด ์ฆ๊ฐ์ ๋ฐ๋ฅธ ์ฑ๋ฅ ๋ณํ๊ฐ ๊ฒ์ฆ๋์ง ์์๋ค.
- ๊ธ์ตยท๋ฒ๋ฅ ์ธ ๋๋ฉ์ธ์ ์ผ๋ฐํ ์ฑ๋ฅ์ด ํ์ธ๋์ง ์์๋ค.
- ๊ตฌ์กฐํ ์ถ๋ ฅ ํ๊ทธ์ ๋๋ฝยท์ค๋ณตยท์์ ์ค๋ฅ์ ๋ํ ํ์ฑ ์คํจ์จ์ด ์ธก์ ๋์ง ์์๋ค.
- ๋ฒ ์ด์ค ๋ชจ๋ธ์ ๋ฉํฐ๋ชจ๋ฌ์ด์ง๋ง ๋ณธ evaluator์ ์ด๋ฏธ์ง ์ ๋ ฅ ์ฑ๋ฅ์ ๊ฒ์ฆ๋์ง ์์๋ค.
7. ๋ผ์ด์ ์ค
7.1 ์ ์ฉ ๋ผ์ด์ ์ค ๊ตฌ์กฐ
๋ณธ ๋ชจ๋ธ์๋ ๋ค์ ์กฐ๊ฑด์ด ํจ๊ป ์ ์ฉ๋๋ค.
- Google์ Gemma Terms of Use
- Selectstar๊ฐ ์์ฒด ์์ ๋ถ์ ๋ถ์ฌํ CC BY-NC 4.0 ๊ธฐ๋ฐ ์ถ๊ฐ ๋น์์ ์กฐ๊ฑด
| ๊ตฌ์ฑ ์์ | ์ ์ฉ ์กฐ๊ฑด |
|---|---|
๋ฒ ์ด์ค ๋ชจ๋ธ google/gemma-3-12b-it |
Gemma Terms of Use |
| Gemma ๊ธฐ๋ฐ ํ์ ๊ฐ์ค์น | Gemma Terms of Use |
| Selectstar๊ฐ ๋ณด์ ํ ์์ ๋ถ | CC BY-NC 4.0 ๊ธฐ๋ฐ ์ถ๊ฐ ๋น์์ ์กฐ๊ฑด |
| ํ์ต ๋ฐ์ดํฐ K2-Feedback | ์๋ณธ ๋ฐ์ดํฐ์ ๋ผ์ด์ ์ค ํ์ธ ํ์ |
| ํ์ตยทํ๊ฐ ์ฝ๋ | ์ ์ฅ์๋ณ ๋ผ์ด์ ์ค ํ์ธ ํ์ |
Hugging Face YAML์์๋ ๋จ์ผ ๋ผ์ด์ ์ค๋ก ์คํด๋์ง ์๋๋ก license: other๋ก ํ๊ธฐํ์๋ค.
7.2 ํ์ฉ๋๋ ์ด์ฉ
- ๋น์์ ์ ํ์ ์ฐ๊ตฌ
- ๋น์๋ฆฌ ๊ต์ก
- ๊ฐ์ธ ํ์ต ๋ฐ ์คํ
- ์ถ์ฒ์ ๋ณ๊ฒฝ์ฌํญ์ ๋ช ์ํ ๋น์์ ์ ์์ ยท์ฐ๊ตฌ
์ฌ๋ฐฐํฌ๋ Gemma Terms์ Selectstar์ ์ถ๊ฐ ์กฐ๊ฑด์ ๋ชจ๋ ์ถฉ์กฑํด์ผ ํ๋ฏ๋ก ๋ณ๋ ๊ฒํ ๊ฐ ํ์ํ๋ค.
7.3 ๋ณ๋ ๊ณ์ฝ ์์ด ํ์ฉ๋์ง ์๋ ์ด์ฉ
- ์ ๋ฃ ์ ํยท์๋น์ค์ ๋ชจ๋ธ ํตํฉ
- ์ ๋ฃ API ๋๋ hosted service ์ ๊ณต
- ์ฌ๋ด ์์ ์ด์ ์์คํ ์ด๋ ๊ณ ๊ฐ ๋๋ฉด ์๋น์ค ์ ์ฉ
- ๋ชจ๋ธ ๊ฐ์ค์น์ ํ๋งค ๋๋ ์์ ์ ์ฌ๋ฐฐํฌ
- ์์ ์ ๋ชจ๋ธ ํ์ต์ ์ํ ๋ฐ์ดํฐ ์์ฑ
- ๊ธฐํ ์ง์ ์ ยท๊ฐ์ ์ ์์ต ์ฐฝ์ถ ๋ชฉ์ ์ ์ด์ฉ
์์ ์ ์ด์ฉ์๋ Selectstar์ ๋ณ๋ ๊ณ์ฝ์ด ํ์ํ๋ฉฐ, ๋ณ๋ ๊ณ์ฝ ํ์๋ Gemma Terms๋ ๊ณ์ ์ ์ฉ๋๋ค.
7.4 Gemma ํ์ ๋ชจ๋ธ ๋ฐฐํฌ ์กฐ๊ฑด
Gemma ํ์ ๋ชจ๋ธ์ ์ 3์ ์ ๊ณต๊ณผ hosted service๋ Gemma Terms์ Distribution์ ํฌํจ๋ ์ ์๋ค. ๋ฐฐํฌ ์์๋ ๋ค์ ์กฐ๊ฑด์ ํ์ธํด์ผ ํ๋ค.
- Gemma ์ฌ์ฉ ์ ํ์ ์งํ ๊ฐ๋ฅํ ์กฐ๊ฑด์ผ๋ก ํฌํจ
- ์๋ น์ธ์๊ฒ Gemma Terms ์ ์ฉ ์ฌ์ค ๊ณ ์ง ๋ฐ ์ฝ๊ด ์ฌ๋ณธ ์ ๊ณต
- ์์ ๋ ํ์ผ์ ๋ณ๊ฒฝ ์ฌ์ค ํ์
- ์๊ตฌ๋๋
NOTICEํ์ผ ์ ๊ณต - Gemma Prohibited Use Policy ์ค์
- ์ถ๊ฐ ์กฐ๊ฑด์ด Gemma Terms์ ์ถฉ๋ํ์ง ์๋์ง ํ์ธ
7.5 ํ์ต ๋ฐ์ดํฐ ๋ผ์ด์ ์ค
K2-Feedback์ ์ ํํ ๋ผ์ด์ ์ค๋ช ๊ณผ ํ์ ๋ฐ์ดํฐ ๋ฐ ๋ชจ๋ธ ํ์ต ํ์ฉ ๋ฒ์๊ฐ ํ์ฌ ๋ฌธ์์ ๊ธฐ์ฌ๋์ด ์์ง ์๋ค. ํ์ธ ์ ๊น์ง ํ์ต ๋ฐ์ดํฐ ์๋ฌธ์ ๋ชจ๋ธ ์ ์ฅ์์ ํฌํจํ๊ฑฐ๋ ์ธ๋ถ์ ์ฌ๋ฐฐํฌํ์ง ์๋๋ค.
7.6 ๊ถ์ฅ ๋ผ์ด์ ์ค ๊ณ ์ง
E-Star-12B-v2-Base is a model derivative of google/gemma-3-12b-it.
Use and distribution of the model weights are subject to the Gemma
Terms of Use. Additional non-commercial terms based on CC BY-NC 4.0
apply to modifications owned by Selectstar.
Users must comply with both sets of terms. Commercial use requires
a separate agreement with Selectstar and remains subject to the
Gemma Terms of Use and the Gemma Prohibited Use Policy.
7.7 ์์ ์ ์ด์ฉ ๋ฌธ์
์์ ์ ๋ผ์ด์ ์ค๊ฐ ํ์ํ ๊ฒฝ์ฐ Selectstar ๊ณต์ ์ฑ๋์ ํตํด ๋ฌธ์ํ๋ค.
- Downloads last month
- 279