Streaming ST
Collection
Streaming speech translation models and frameworks • 3 items • Updated
How to use pltobing/translategemma-4b-it-Q8_0-GGUF with Transformers:
# Use a pipeline as a high-level helper
# Warning: Pipeline type "translation" is no longer supported in transformers v5.
# You must load the model directly (see below) or downgrade to v4.x with:
# pip install "transformers<5.0.0"
from transformers import pipeline
pipe = pipeline("translation", model="pltobing/translategemma-4b-it-Q8_0-GGUF") # pip install -U transformers accelerate
# Load model directly
from transformers import AutoModel
model = AutoModel.from_pretrained("pltobing/translategemma-4b-it-Q8_0-GGUF", device_map="auto")How to use pltobing/translategemma-4b-it-Q8_0-GGUF with llama.cpp:
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf pltobing/translategemma-4b-it-Q8_0-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf pltobing/translategemma-4b-it-Q8_0-GGUF:Q8_0
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf pltobing/translategemma-4b-it-Q8_0-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf pltobing/translategemma-4b-it-Q8_0-GGUF:Q8_0
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf pltobing/translategemma-4b-it-Q8_0-GGUF:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf pltobing/translategemma-4b-it-Q8_0-GGUF:Q8_0
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf pltobing/translategemma-4b-it-Q8_0-GGUF:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf pltobing/translategemma-4b-it-Q8_0-GGUF:Q8_0
docker model run hf.co/pltobing/translategemma-4b-it-Q8_0-GGUF:Q8_0
How to use pltobing/translategemma-4b-it-Q8_0-GGUF with Ollama:
ollama run hf.co/pltobing/translategemma-4b-it-Q8_0-GGUF:Q8_0
How to use pltobing/translategemma-4b-it-Q8_0-GGUF with Docker Model Runner:
docker model run hf.co/pltobing/translategemma-4b-it-Q8_0-GGUF:Q8_0
How to use pltobing/translategemma-4b-it-Q8_0-GGUF with Lemonade:
# Download Lemonade from https://lemonade-server.ai/ lemonade pull pltobing/translategemma-4b-it-Q8_0-GGUF:Q8_0
lemonade run user.translategemma-4b-it-Q8_0-GGUF-Q8_0
lemonade list
This repository is publicly accessible, but you have to accept the conditions to access its files and content.
To access Gemma on Hugging Face, you’re required to review and agree to Google’s usage license. To do this, please ensure you’re logged in to Hugging Face and click below. Requests are processed immediately.
Log in or Sign Up to review the conditions and access this model content.
8-bit
Base model
google/translategemma-4b-it