Instructions to use Qwen/Qwen3.8-27B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Qwen/Qwen3.8-27B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Qwen/Qwen3.8-27B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Qwen/Qwen3.8-27B") model = AutoModelForMultimodalLM.from_pretrained("Qwen/Qwen3.8-27B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Inference
- HuggingChat
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Qwen/Qwen3.8-27B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Qwen/Qwen3.8-27B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Qwen/Qwen3.8-27B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Qwen/Qwen3.8-27B
- SGLang
How to use Qwen/Qwen3.8-27B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Qwen/Qwen3.8-27B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Qwen/Qwen3.8-27B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Qwen/Qwen3.8-27B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Qwen/Qwen3.8-27B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use Qwen/Qwen3.8-27B with Docker Model Runner:
docker model run hf.co/Qwen/Qwen3.8-27B
Beta testers wanted (Qwen 3.8:27b)
Anyone interested in trying out a new inference provider?
We're testing Qwen 3.8:27B behind an OpenAI-compatible API and looking for a few users to give it a spin and tell us how it works for them. You can use it for any regular Qwen 3.8 purpose via any standard client (VSCode, OpenCode, Cline, Pi, etc.) and we will provide free credits. We're looking for feedback, bug reports, etc. in return.
To anticipate a few questions: 1) Yes we are "Zero Data Retention" (formal ZDR policy document is in the works), 2) Our offering is centered around some core technology that can compile models down to a very efficient executable allowing us to deploy (and collocate them) very quickly and flexibly, and 3) We are being a bit coy about the actual identity of the company because we are still in (semi-)stealth mode.
If you're interested, please shoot me an email at anon@m00se.com (I'm also happy to answer questions or provide more details here). Thanks!
Have you considered alternative models like Swift-Qwen3.8-27b?
Yes indeed, we are looking at deploying more models in this area, but have started with just a handful for initial beta testing.
Was there a very particular feature of Swift-Qwen3.8-27b that you were really interested in?
Also, do you have any sense of what sort of token volume you'd be interested in over, say, per month?
I'm asking because it might well be possible to prioritize getting Swift-Qwen3.8-27b implemented, but interested in a sense of potential use volume.
Thanks!
We may indeed have some interest in helping devs fine tune their models, so anything you can share about your interest in SwiftQwen in particular would be welcome.
Im interested in knowing what data protection you have. You must know by now that people who run local models care a lot about their data, where its being sent and stored.
Are our prompts being saved in some way, do the endpoints have some form of encryption against possible interception attacks on your end? And what would be the inference speed you allocate for users (as in tok/s infill and decode)?
Yes, very good point. Should have mentioned privacy and data retention in my own post.
We do have a formal ZDR policy which is pretty standard, with basically an exception only for incidental logging (e.g. an error report of a system failure will sometimes included fragments of the request or response that may have been in process when the failure occurred).
Our raison d'etre is trying to compile and run existing models as efficienly as possible, not to do anything with user data. There is standard (SSL) encryption to the endpoint and ZDR in process and response. Exact TPS is part of the beta testing process, but certainly 200 and above as a target (the beta testing models are generally running on B200 hardware), but time-to-first token is a larger question mark depending on which ones are loaded and which need a cold start (but our "cold" starts are also generally really fast, and again, part of the beta testing process to prove that out).
And what about context lenght? Ive noticed a significant drop off of quality in code after the 500k token mark on qwen models at f16 kv, are users going to be able to choose their context lenght and the quantization?
I say this because im not sure qwens 1M context is actually fair for any quant except bf16 (though i havent tested it, i dont have the vram for it). So users might be interested in running smaller contexts if they can pay less for the capacity.
Also you didnt mention the pricing per million token. What pricing can we expect from qwen 27b or qwen flash next, even the newest glm flash if you have capacity for it.
We don't (yet) support YaRN, so our Qwen-3.8 context length is currently the model's native context length (i.e. 262,144). All of our tests so far have been based on Nvidia FP4 fr quantization, but that is not set in stone by any means.
No pricing yet; this is a real product beta testing, a large part of which is to prove out if we can do this at scale with custom models in a sustainable way. We would provide free usage for beta testers while we figure out what we would even charge for this. That should become apparent from the level of real world efficiency we can get from the test (e.g. if we can get 3 custom models running at x Tok/s on a B200, we take the hourly cost of the B200 and divide across all those x token stream to try and figure out a sustainable token pricing level).