Instructions to use gaokerena/amestris-1b-sft with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use gaokerena/amestris-1b-sft with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Amestris-1B-SFT
amestris-1b-sft is the supervised fine-tuning (SFT) comparison checkpoint from the Amestris English-to-German machine-translation study. It adapts google/gemma-3-1b-it with LoRA using approximately 27,000 difficult prompt–reference pairs. The rejected_answer field retained in the research dataset was used for auditing and comparison, not in the SFT loss.
This repository contains a non-merged PEFT/LoRA adapter, not a standalone model. In this release, the loadable adapter is stored in the sft_hard27k_lora_adapter/ subfolder; the subfolder argument in the example below is therefore required.
Model details
| Field | Value |
|---|---|
| Task | English → German machine translation |
| Architecture | Gemma 3 1B instruction-tuned causal language model + LoRA |
| Base model | google/gemma-3-1b-it |
| Post-training objective | Supervised next-token prediction on prompt → preferred translation |
| Adapter location | sft_hard27k_lora_adapter/ |
| Adapter type | PEFT LoRA, non-merged |
| LoRA rank / alpha / dropout | 32 / 32 / 0.05 |
| Maximum sequence length | 768 tokens |
| Training epochs | 1 |
| Learning rate | 5 × 10⁻⁵ |
| Random seed | 42 |
| Languages | English input, German output |
The wider research methodology and comparison with DPO are documented in the project repository and the associated paper.
Quick start: run the model
1. Install dependencies
pip install -U "transformers>=4.50.0" "peft>=0.18.1" "accelerate>=1.0" torch
Accept Google’s Gemma terms on the base-model page, then authenticate with an account that has access:
hf auth login
2. Load the adapter and translate
import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
BASE_MODEL_ID = "google/gemma-3-1b-it"
ADAPTER_ID = "gaokerena/amestris-1b-sft"
ADAPTER_SUBFOLDER = "sft_hard27k_lora_adapter"
# Use the gated base-model tokenizer. This is also the tokenizer source used
# by the project’s evaluation workflow.
tokenizer = AutoTokenizer.from_pretrained(BASE_MODEL_ID)
base_model = AutoModelForCausalLM.from_pretrained(
BASE_MODEL_ID,
device_map="auto",
dtype="auto",
)
model = PeftModel.from_pretrained(
base_model,
ADAPTER_ID,
subfolder=ADAPTER_SUBFOLDER,
)
model.eval()
source_text = "The committee will publish its final report next Tuesday."
translation_prompt = f"""You are an expert professional translator from English to German.
Preserve all meaning, nuance, factual details, names, dates, and numbers.
Use natural, fluent, professional standard German.
Do not omit, summarize, add, or explain anything.
Return only the German translation.
Text to translate:
{source_text}"""
messages = [{"role": "user", "content": translation_prompt}]
formatted_prompt = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
)
inputs = tokenizer(formatted_prompt, return_tensors="pt")
device = next(model.parameters()).device
inputs = {name: tensor.to(device) for name, tensor in inputs.items()}
with torch.inference_mode():
output_ids = model.generate(
**inputs,
max_new_tokens=256,
do_sample=False,
num_beams=4,
early_stopping=True,
repetition_penalty=1.05,
no_repeat_ngram_size=3,
)
generated_ids = output_ids[0, inputs["input_ids"].shape[1]:]
translation = tokenizer.decode(generated_ids, skip_special_tokens=True).strip()
print(translation)
The subfolder parameter is specific to this repository layout. Omitting it causes PEFT to search the repository root for adapter_config.json, where this release does not store the SFT adapter.
Training methodology
For SFT, each training example maps the standardized English translation prompt to prefered_answer (the spelling used in the project artifact). The rejected_answer column was retained only for auditability and was not included in the supervised objective. Training used a one-epoch LoRA configuration with rank 32, alpha 32, dropout 0.05, maximum sequence length 768, learning rate 5 × 10⁻⁵, and seed 42.
The saved training metadata reports a final training loss of 1.7378 after 1,570 optimization steps, approximately 7.95 samples per second, and 7,825,728 input tokens processed. These values document the recorded run; they should not be treated as cross-system performance benchmarks.
Evaluation
The following English-to-German WMT14 results are reported by the project
| Metric | gemma-3-1b baseline | amestris-1b-sft | better direction |
|---|---|---|---|
| BLEU | 0.1573 | 0.1701 | ↑ |
| COMET22 | 0.7698 | 0.8007 | ↑ |
| COMET-KIWI22 | 0.7031 | 0.7738 | ↑ |
| METEOR | 0.3861 | 0.4297 | ↑ |
| TER | 0.7765 | 0.7840 | ↓ |
| chrF++ | 41.93 | 44.87 | ↑ |
These values are reported by the project and were not independently reproduced during preparation of this card. Fair comparison requires identical prompts, decoding, preprocessing, test examples, and metric versions.
Intended use
This checkpoint is intended for:
- English-to-German translation research;
- controlled comparison of SFT and preference-based post-training;
- PEFT/LoRA ablation studies on a compact causal language model;
- reproducibility studies using the Amestris hard-27k data configuration.
It is not certified for legal, medical, financial, safety-critical, or unsupervised production translation.
Limitations and risks
- The model is specialized for English-to-German translation; other languages and directions are out of scope.
- The compact base model can struggle with long context, specialized terminology, rare entities, and structurally complex sentences.
- Automatic metrics do not guarantee factual or terminological correctness.
- Names, numbers, dates, formatting, or source details may be altered or omitted; sensitive translations require human verification.
- Biases and safety limitations from the base model and training data may remain.
- Use is governed by the Gemma license and base-model access conditions.
Repository contents
sft_hard27k_lora_adapter/: final PEFT adapter, tokenizer artifacts, training metrics, and a trainer checkpoint.metadata/: upload and reproducibility metadata.archives/: archived checkpoint package and checksum.README.md: this model card.
The nested trainer checkpoint and archive are provided for research reproducibility. Standard inference should load the final adapter subfolder shown in the quick-start example.
Citation
@misc{ghassabi2026backtranslation,
title = {Backtranslation Augmented Direct Preference Optimization for Neural Machine Translation},
author = {Ghassabi, Mehrdad and Rajabi, Spehr and Baradaran Kashani, Hamidreza and Hakim, Sadra and Keivandarian, Mahshid and Jahani Bahnamiri, Amirhossein},
year = {2026},
eprint = {2604.25702},
archivePrefix = {arXiv},
primaryClass = {cs.CL}
}
Acknowledgments
This checkpoint was developed as part of the Amestris research project by Mehrdad Ghassabi, Sepehr Rajabi, Hamidreza Baradaran Kashani, Sadra Hakim, Mahshid Keivandarian, and Amirhossein Jahani Bahnamiri.
- Downloads last month
- -