SmolVLM-500M-OSAgent — sft_r256_step6500

Author: Gabriel Schwarzbauer

A compact 507M-parameter multimodal vision-language model fine-tuned to autonomously operate a simulated Windows 11 desktop. It observes screen pixels (1280×800 viewport), user instructions, and action history, and outputs one structured mouse or keyboard action per step (e.g. click 14 977).

The model is a fine-tune of HuggingFaceTB/SmolVLM-500M-Instruct (~507M parameters, bfloat16, ~969 MB standalone weights) using LoRA (rank 256, alpha 512) merged directly into base weights.

Independent Academic Research — No Affiliation with Microsoft:
This open-source research artifact is developed independently by Gabriel Schwarzbauer. It is not affiliated with, associated with, sponsored by, endorsed by, authorized by, or in any way officially connected with Microsoft Corporation, or any of its subsidiaries or affiliates.

Read this first (Research Disclaimer & Watermark Notice):

  • Active Research Checkpoint: This release evaluates Level 0 (Shell Primitives) and Level 1 (File Explorer, Notepad & Desktop Management). On our standardized closed-loop benchmark (410 tasks across 41 families, 10 variations per family), it achieves 93.3% success on Level 0 and 49.4% on Level 1 (59.0% overall, average trajectory score 0.84). While strong in core navigation, window management, and file interactions, it still fails certain inline editing tasks and is not a finished commercial product.
  • Zero Desktop Duplication (Full Randomization): The agent never saw the same desktop state twice during training or evaluation. Every single episode operates on a freshly seeded, procedurally randomized desktop environment (random themes, accent colors, wallpapers, display DPI scaling, pinned applications, file clutter, Wi-Fi networks, Bluetooth devices, and window geometries).
  • Icon Watermark & Copyright Protection: A fine 16px black coordinate grid with subtle alternating dark box shading is applied across all demonstration screenshots. This watermarks and protects third-party application icons and trademarks against unauthorized reproduction while preserving overall visual legibility of desktop layouts, UI controls, and agent interactions.
  • Simulator vs. Real Windows 11 Transfer: The model was trained and quantitatively benchmarked inside a deterministic web-based simulation of a Windows 11 desktop (1280×800). While preliminary exploratory live tests on physical Windows 11 installations have demonstrated highly positive zero-shot transfer (autonomously navigating native Start menus, taskbars, and File Explorer), these physical system tests are not part of the standardized benchmark statistics reported above. When deploying on native hardware, variations in display scaling, system font rendering, and custom OS themes may affect coordinate precision.
  • Safety & Isolation: The model contains no internal safety filters or permission systems. Do not connect this model to a real desktop or production environment without strict sandbox isolation.
  • Independence & Trademarks: This project is completely independent of Microsoft Corporation. "Microsoft", "Windows", "Windows 11", "File Explorer", and "Notepad" are registered trademarks of Microsoft Corporation. All third-party application names, trademarks, and logos (e.g. LibreOffice, Brave, Wireshark, NVIDIA, Git, Steam) belong strictly to their respective owners. References to these names and interfaces are made solely for nominative, educational, and scientific identification within our simulated research environment.

Zero-Duplication Procedural Desktop Randomization

To prevent memorization of static pixel coordinates or visual cues, the simulation implements comprehensive procedural randomization. Every episode receives a unique random seed — the agent never operated on the same desktop twice in training or evaluation:

  1. Appearance & Themes:
    • Color Mode: Randomly drawn Dark Mode vs. Light Mode (50% / 50%).
    • Wallpapers: Randomly assigned from official Windows 11 packs (Windows Dark/Light, Glow 1–4, Captured Motion 1–4, Sunrise 1–4, Flow 1–4) or solid color backgrounds (25% probability).
    • Accent Colors: Randomly selected from 48 authentic Windows accent palette swatches (e.g., teal #038387, dark blue #0063b1, purple #744da9, yellow gold #ffb900).
  2. Display Scaling (DPI):
    • Dynamic scale factors (125%, 150%, and 175%) that alter font sizes, button dimensions, and padding.
  3. Taskbar & Start Menu Pinned Applications:
    • Start Menu: A random sample of 6 to 12 apps drawn from the full application registry.
    • Taskbar: A random sample of 1 to 4 apps pinned dynamically to the taskbar.
  4. Desktop & Filesystem Clutter:
    • Application Shortcuts: Random desktop shortcuts with real application icons (LibreOffice, NVIDIA App, Git Bash, Raspberry Pi Imager, CrystalDiskMark, etc.).
    • Folder Hierarchies: Randomly named directories and subfolders (up to 2 levels deep) with realistic directory names drawn from vocabulary word pools.
    • Files: Procedurally generated files with varying extensions (.txt, .docx, .pdf, .xlsx, .png, .jpg) and randomized creation/modification timestamps.
  5. Wireless & Connectivity State:
    • Wi-Fi: Randomly generated lists of 3 to 6 Wi-Fi SSIDs with realistic naming roots and suffixes (-5G, _Home, WiFi, 2.4GHz, _Guest), varied signal strengths (1–4 bars), and security statuses.
    • Bluetooth: Randomized Bluetooth toggle states and discoverable device names.
  6. Task Manager & System Load:
    • Random subset of background processes (40% to 100% of active services).
    • Randomized initial process table sorting columns (name, cpu, memory, disk, network, gpu) and directions (asc / desc).
    • Dynamic CPU and memory load multipliers (0.4x to 2.0x).
  7. Window Geometries & Layouts:
    • Every application window opens at a randomized width and height (0.7x to 1.3x base dimensions) and dynamic position offsets, ensuring the model must visually locate buttons and tabs rather than hardcoding coordinates.
  8. File Explorer View & Sort Settings:
    • Randomized initial view modes (xl, large, medium, small, list, details, tiles, content).
    • Randomized sort keys (name, modified, type, size) and sort directions (asc / desc).

Step-by-Step Trajectory Walkthroughs (Flawless Execution Examples)

The following real closed-loop episodes illustrate the step-by-step visual decision making of the model on complex, multi-stage Level 1 tasks. For every step, the model observes the current screen pixels, instruction, and action history, and outputs the exact next action.

(A fine 16px black coordinate grid with subtle alternating dark box shading is overlaid across all images to watermark proprietary application icons and protect copyright.)

Example 1: Open the C: Drive in File Explorer (explorer_open_drive)

  • Task Prompt: "Open the C: drive in File Explorer."
  • Random Seed: 1614522484 (Randomized desktop: custom wallpaper, dark theme, custom pinned apps, dynamic window positioning)
  • Outcome: ⭐ 100% PERFECT SUCCESS (Executed in 5 actions + done)
Step 0: Desktop with Procedural Clutter Step 1: Start Menu Opened Step 2: Search Query Entered
Step 0 Step 1 Step 2
Model Action: click 15 974
(Clicks Start button on taskbar)
Model Action: type "file explorer"
(Types app name into Start search)
Model Action: key enter
(Launches File Explorer window)
Step 3: File Explorer Window Open Step 4: Drives Displayed in This PC Step 5: Inside C: Drive (Complete)
Step 3 Step 4 Step 5
Model Action: click 310 484
(Clicks "This PC" in left navigation pane)
Model Action: double_click 454 459
(Double-clicks C: drive to open)
Model Action: done
(C: drive content displayed, task complete)

Example 2: Open a New Tab in Notepad (notepad_new_tab)

  • Task Prompt: "In Notepad, open a new tab."
  • Random Seed: 1681262377 (Randomized desktop: custom wallpaper, dark theme, custom taskbar search, dynamic window layout)
  • Outcome: ⭐ 100% PERFECT SUCCESS (Executed in 5 actions + done)
Step 0: Desktop with Procedural Clutter Step 1: Taskbar Search Flyout Step 2: App Search Result
Step 0 Step 1 Step 2
Model Action: click 118 978
(Clicks taskbar search bar)
Model Action: type "notepad"
(Types app name into search)
Model Action: key enter
(Launches Notepad application)
Step 3: Notepad Application Open Step 4: File Menu Opened Step 5: Second Tab Created (Complete)
Step 3 Step 4 Step 5
Model Action: click 217 338
(Clicks "File" menu in Notepad)
Model Action: click 297 377
(Clicks "New tab" in menu)
Model Action: done
(Two tabs active, task complete)

Action Space & Grammar

The model generates exactly one action string per step matching this grammar:

Action Syntax Meaning
Click click X Y Left mouse click at normalized coordinates
Double Click double_click X Y Double click (e.g. to open desktop icon or file)
Right Click right_click X Y Context menu click
Type type "text" Types string (\n = Enter, escaped quotes)
Key key <key> Single key or hotkey combo (key enter, key ctrl+s, key win)
Scroll scroll X Y N Mouse wheel scroll by N ticks (positive = down)
Drag drag X1 Y1 X2 Y2 Click and drag from (X1, Y1) to (X2, Y2)
Wait wait Pause for animations or loading state
Done done Declares the task successfully completed
  • X, Y are normalized integers from 0 to 999: (0, 0) is top-left, (999, 999) is bottom-right.
  • Pixel coordinate conversion: pixel_x = (X + 0.5) / 1000 * width, pixel_y = (Y + 0.5) / 1000 * height.

Prompt Format

Task: <task description>
Actions so far: <"none" or "1) click 14 977; 2) type \"report\"; ...">
Next action:

Quickstart / Inference

import re
import torch
from PIL import Image
from transformers import AutoProcessor, AutoModelForImageTextToText

repo = "PATH_OR_HF_REPO_ID"          # Path to merged model folder or Hugging Face Hub ID
device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.bfloat16 if device == "cuda" else torch.float32

processor = AutoProcessor.from_pretrained(repo)
processor.image_processor.size = {"longest_edge": 1536}
processor.image_processor.max_image_size = {"longest_edge": 512}
model = AutoModelForImageTextToText.from_pretrained(repo, dtype=dtype).to(device).eval()

eos = [processor.tokenizer.convert_tokens_to_ids("<end_of_utterance>"), processor.tokenizer.eos_token_id]

def next_action(screenshot: Image.Image, task: str, past_actions: list[str]) -> str:
    screenshot = screenshot.convert("RGB").resize((1280, 800))
    shown = past_actions[-12:]
    start = len(past_actions) - len(shown)
    history = "; ".join(f"{start + i + 1}) {a}" for i, a in enumerate(shown)) or "none"
    prompt = f"Task: {task}\nActions so far: {history}\nNext action:"
    
    messages = [{"role": "user", "content": [{"type": "image"}, {"type": "text", "text": prompt}]}]
    text = processor.apply_chat_template(messages, add_generation_prompt=True)
    inputs = processor(text=[text], images=[[screenshot]], return_tensors="pt").to(device)
    inputs["pixel_values"] = inputs["pixel_values"].to(dtype)
    
    with torch.inference_mode():
        out = model.generate(**inputs, max_new_tokens=28, do_sample=False, eos_token_id=eos)
    answer = processor.tokenizer.decode(out[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True)
    return answer.strip().split("\n")[0].strip()

# Example test
history = []
img = Image.new("RGB", (1280, 800), color=(30, 30, 30))
action = next_action(img, "Show the Start menu.", history)
print("Predicted action:", action)      # e.g. click 19 984

Benchmark Results (Honest Evaluation)

All metrics were evaluated in closed-loop interactive simulations with 10 distinct randomized hold-out variations per task family (total 410 evaluation episodes across Level 0 and Level 1). Held-out words, directories, and file names were strictly isolated (split="eval"), ensuring the agent never saw these parameters during training.

  • Evaluation Settings: Temperature T = 0.3 (slight sampling variation), step budget = 3.0x Demonstration.
  • Success: Binary verification that all target operating system states and events occurred in correct order.
  • Score: Weighted Longest Common Subsequence (Weighted LCS) of reached milestone events (reflecting partial trajectory progress).

1. High-Level Summary by Level

Level Families Tested Variations Total Tasks Success Rate (%) Average Score
Level 0 (Shell Primitives) 9 of 9 10 90 93.3% (84/90) 0.96
Level 1 (Explorer & Desktop Management) 32 families 10 320 49.4% (158/320) 0.81
COMBINED TOTAL 41 families 10 410 59.0% (242/410) 0.84

2. Detailed Results by Application

run — 100.0% (10/10), score 1.00 — 1 task families

Family Task Description Level Success Rate Avg Score Status
run_cancel Open Run and cancel it Level 0 100.0% 1.00 ⭐ 100% PERFECT

notepad — 76.7% (23/30), score 0.91 — 3 task families

Family Task Description Level Success Rate Avg Score Status
notepad_new_tab Open a second tab Level 1 100.0% 1.00 ⭐ 100% PERFECT
notepad_open_file Open a file with File > Open Level 1 100.0% 0.99 ⭐ 100% PERFECT
notepad_write_save Write a note and save it Level 1 30.0% 0.75 🟡 PARTIAL SUCCESS

shell — 67.4% (128/190), score 0.87 — 19 task families

Family Task Description Level Success Rate Avg Score Status
open_task_manager_menu Open Task Manager from a menu Level 0 100.0% 1.00 ⭐ 100% PERFECT
winx_menu_pages Win+X menu shortcut to a Settings page Level 0 100.0% 1.00 ⭐ 100% PERFECT
unpin_from_taskbar Unpin File Explorer from the taskbar Level 0 100.0% 1.00 ⭐ 100% PERFECT
open_flyout Open a taskbar flyout Level 0 100.0% 1.00 ⭐ 100% PERFECT
open_start_menu Open the Start menu Level 0 100.0% 1.00 ⭐ 100% PERFECT
open_recycle_bin Open the Recycle Bin Level 0 100.0% 1.00 ⭐ 100% PERFECT
open_app Open an app Level 1 100.0% 1.00 ⭐ 100% PERFECT
pin_to_taskbar Pin an app to the taskbar Level 1 90.0% 0.97 🟢 VERY GOOD
quick_toggle Toggle a quick setting Level 0 80.0% 0.90 🟢 VERY GOOD
window_snap Snap a window to a screen half Level 1 60.0% 0.95 🟡 PARTIAL SUCCESS
desktop_icon_size Change the desktop icon size Level 1 60.0% 0.80 🟡 PARTIAL SUCCESS
desktop_settings_link Desktop menu: Display settings / Personalize Level 0 60.0% 0.73 🟡 PARTIAL SUCCESS
desktop_new_text New text document on the desktop Level 1 50.0% 0.67 🟡 PARTIAL SUCCESS
window_move Move a window Level 1 40.0% 0.84 🟡 PARTIAL SUCCESS
unpin_from_start Unpin an app from Start Level 1 40.0% 0.73 🟡 PARTIAL SUCCESS
desktop_new_folder New folder on the desktop Level 1 40.0% 0.60 🟡 PARTIAL SUCCESS
window_resize Resize a window by an edge Level 1 20.0% 0.90 🔴 CHALLENGING
window_op Maximize / minimize / close a window Level 1 20.0% 0.83 🔴 CHALLENGING
desktop_sort Sort the desktop icons Level 1 20.0% 0.60 🔴 CHALLENGING

file-explorer — 45.0% (81/180), score 0.79 — 18 task families

Family Task Description Level Success Rate Avg Score Status
explorer_open_this_pc Open This PC Level 1 100.0% 1.00 ⭐ 100% PERFECT
explorer_open_drive Open the C: drive Level 1 100.0% 1.00 ⭐ 100% PERFECT
explorer_details_button Toggle the details pane Level 1 100.0% 1.00 ⭐ 100% PERFECT
explorer_properties Open Properties of a file Level 1 90.0% 0.97 🟢 VERY GOOD
explorer_show_option Turn on a View > Show option Level 1 80.0% 0.95 🟢 VERY GOOD
explorer_open_file Open a text file Level 1 70.0% 0.85 🟢 VERY GOOD
explorer_select_all Select all items Level 1 60.0% 0.90 🟡 PARTIAL SUCCESS
explorer_go_up Go up one level Level 1 50.0% 0.84 🟡 PARTIAL SUCCESS
explorer_delete Delete a file Level 1 50.0% 0.75 🟡 PARTIAL SUCCESS
explorer_open_folder Open a folder Level 1 40.0% 0.87 🟡 PARTIAL SUCCESS
explorer_sort Sort the file list Level 1 30.0% 0.78 🟡 PARTIAL SUCCESS
explorer_move_file Move a file to another folder Level 1 20.0% 0.69 🔴 CHALLENGING
explorer_copy_file Copy a file to another folder Level 1 10.0% 0.69 🔴 CHALLENGING
explorer_view Change the folder view Level 1 10.0% 0.59 🔴 CHALLENGING
explorer_search Search in a folder Level 1 0.0% 0.69 🔴 CHALLENGING
explorer_rename Rename a file Level 1 0.0% 0.57 🔴 CHALLENGING
explorer_new_folder Create a folder Level 1 0.0% 0.53 🔴 CHALLENGING
explorer_new_text Create a text document Level 1 0.0% 0.50 🔴 CHALLENGING

Strengths, Breakthroughs & Known Limitations

Where the Model Excels:

  1. Level 0 Shell Primitives (Virtually Mastered):
    • 7 out of 9 Level 0 families scored ⭐ 100% PERFECT across all 10 variations (open_start_menu, open_task_manager_menu, winx_menu_pages, unpin_from_taskbar, open_flyout, open_recycle_bin, run_cancel).
    • Quick settings toggling achieved 80.0% (score 0.90).
  2. Breakthrough on Previously Unsolved Tasks:
    • explorer_move_file: Achieved 20.0% success (score 0.69).
    • desktop_sort: Achieved 20.0% success (score 0.60).
    • explorer_view: First successful completions (10.0%, score 0.59).
  3. High Trajectory Alignment:
    • An overall average score of 0.84 confirms that even in incomplete episodes, the model reliably executes the majority of intermediate steps (opening windows, navigating to directories, selecting files) before halting.
  4. Editor and Window Mechanics:
    • Notepad tab operations and opening files achieved 100.0% and 99.0% success.
    • Window movement and window snapping to screen halves operate with high fidelity (scores 0.84 – 0.95).

Transparent Limitations & Areas for Improvement:

  1. In-Place File Explorer Text Entry (The 4 Unsolved Tasks):
    • Only 4 families remain at 0% in Level 1: explorer_search, explorer_rename, explorer_new_folder, and explorer_new_text.
    • Root Cause: All four require inline text editing inside File Explorer list views, where small focus bounding boxes and Enter-key confirmation timing remain difficult for the 500M vision encoder.
  2. Sub-Pixel Edge Dragging (window_resize at 20%):
    • Resizing windows by hovering and dragging the 2-pixel window border is geometrically sensitive.
  3. Temperature Sensitivity:
    • At T = 0.3, the model shows high adaptability and breaks out of click loops, but experiences occasional coordinate jitter on small buttons. For high-precision deterministic runs, T = 0.0 or 0.1 is recommended.

Training Configuration

  • Base Model: HuggingFaceTB/SmolVLM-500M-Instruct
  • Fine-Tuning Method: PEFT LoRA (rank r = 256, alpha = 512, target modules: all linear projections in vision and language backbones)
  • Precision: bfloat16
  • Batch Configuration: Effective batch size 16 (micro-batch 2, gradient accumulation 8)
  • Optimization: AdamW (peak lr = 2e-4, cosine schedule with floor at 7e-5)
  • Curriculum & Sampling: Autonomous dynamic loss-weighted sampling (FamilyWeights), dynamically upweighting harder families while consolidating mastered skills.

Legal & Non-Affiliation Disclaimer

  • Non-Affiliation: This model, research project, and repository are entirely independent and not affiliated with, endorsed by, sponsored by, authorized by, or associated with Microsoft Corporation or any third-party software developers.
  • Nominative Fair Use: References to "Microsoft", "Windows", "Windows 11", "File Explorer", "Notepad", "Task Manager", or third-party application titles are made strictly under nominative fair use for descriptive, academic, and scientific identification of the desktop simulation environment.
  • No Proprietary Microsoft Code or Binaries: This repository does not contain, distribute, or require any proprietary Microsoft Windows binaries, dynamic link libraries (DLLs), system files, or private APIs. All agent interaction occurs within an independent, open-source synthetic web-based simulation environment.
  • Third-Party Trademark Protection: Demonstration screenshots incorporate coordinate grid overlays and cell watermarks to prevent the reproduction or distribution of third-party application icons and graphical trademarks.
Downloads last month
40
Safetensors
Model size
0.5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Gabriel8495839/SmolVLM-500M-OSAgent