Instructions to use Praneyaarora/satquery-internvl35-1b-epoch2-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Praneyaarora/satquery-internvl35-1b-epoch2-lora with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("OpenGVLab/InternVL3_5-1B-Instruct") model = PeftModel.from_pretrained(base_model, "Praneyaarora/satquery-internvl35-1b-epoch2-lora") - Notebooks
- Google Colab
- Kaggle
- SatQuery InternVL3.5-1B Epoch-2 LoRA
- 1. Overview
- 2. Base Model
- 3. Fine-Tuning Method & LoRA Configuration
- 4. Training Dataset (SatQuery 15K SFT Corpus)
- 5. Training Objectives & Architectural Principle
- 6. Training Run Details
- 7. Benchmark Results
- 8. Results Interpretation & Deployment Selection
- 9. Important Limitations
- 10. Exact Coordinate Convention
- 11. Quickstart & Inference
- 12. SatQuery Integration
- 13. License & Compliance
- 14. Citation
- 1. Overview
SatQuery InternVL3.5-1B Epoch-2 LoRA
Official LoRA / PEFT adapter for SatQuery Stage-2 fine-tuning of OpenGVLab/InternVL3_5-1B-Instruct for remote-sensing visual understanding, spatial grounding, and satellite image interpretation.
Note: This repository contains ONLY the LoRA / PEFT adapter weights (
adapter_model.safetensorsandadapter_config.json). The base model weights are not hosted here and must be loaded from OpenGVLab/InternVL3_5-1B-Instruct.
1. Overview
SatQuery is an agentic Earth Observation intelligence system that combines vision-language models with deterministic geospatial analysis tools (e.g. optical/SAR fetchers, vegetation indices, affine projections, terrain analysis, and change detection).
This adapter specializes the compact InternVL3.5-1B-Instruct model for:
- Precise natural-language spatial grounding and tiny-object localization in satellite rasters.
- Grounded visual question answering (VQA) over Earth observation scenes.
- Detailed remote-sensing scene captioning and visual change description.
2. Base Model
| Attribute | Specification |
|---|---|
| Model ID | OpenGVLab/InternVL3_5-1B-Instruct |
| Total Base Parameters | |
| LLM Backbone | Qwen-based Causal LM |
| Vision Encoder | InternVision (~300M parameters, frozen) |
| Projector | MLP Cross-modal Projector (frozen) |
3. Fine-Tuning Method & LoRA Configuration
Fine-tuning was conducted strictly using Parameter-Efficient Fine-Tuning (PEFT / LoRA) on the language model projection layers while keeping the vision encoder and cross-modal projector frozen.
| Parameter | Verified Value |
|---|---|
| PEFT Type | LoRA (LORA) |
| LoRA Rank ($r$) | 16 |
| LoRA Alpha ($\alpha$) | 32 |
| LoRA Dropout | 0.05 |
| Target Modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
| Trainable Parameters | 10,092,544 |
| Total Parameters | 1,070,990,336 |
| Trainable Parameter Fraction | ~0.9424% |
| Precision | BF16 (bfloat16) |
| Gradient Checkpointing | Enabled (true) |
| Micro-batch Size | 1 |
| Gradient Accumulation Steps | 8 |
| Effective Batch Size | 8 |
4. Training Dataset (SatQuery 15K SFT Corpus)
The Stage-2 Supervised Fine-Tuning (SFT) corpus consists of 15,000 curated, deduplicated remote-sensing multimodal instruction pairs:
| Subset | Sample Count | Percentage | Primary Task |
|---|---|---|---|
| Spatial Grounding | 12,000 | 80% | Natural-language bounding-box localization & coordinates |
| Remote Sensing VQA | 2,250 | 15% | Multi-turn reasoning on land cover, infrastructure & water |
| Scene Captioning | 750 | 5% | Dense technical Earth observation descriptions |
| Total Corpus | 15,000 | 100% | 11,326 unique remote sensing images |
- Data Integrity: 0 duplicate records, 0 synthetic oversampling, 0 evaluation leakage.
- Coordinate Convention: All bounding boxes were normalized and formatted strictly as
[ymin, xmin, ymax, xmax]with values in $[0, 1]$ (top-left origin).
5. Training Objectives & Architectural Principle
SFT Objectives:
- Enhance natural-language spatial grounding and tiny-feature localization in aerial/satellite scenes.
- Improve structured coordinate output adherence without requiring external grounding heads.
- Strengthen multi-turn satellite visual conversation capabilities.
Architectural Principle in SatQuery:
The VLM is designed for semantic visual interpretation, NOT for deterministic geospatial mathematics.
In the SatQuery architecture, the VLM handles visual feature extraction and rough image-space grounding. Exact operations (NDVI calculation, CRS reprojecting, affine matrix transformations, polygon area calculations, and meteorological data fetching) are delegated to deterministic, mathematically grounded tools.
6. Training Run Details
- Run ID:
internvl35_satquery_stage2_2epoch_20260927_191016 - Epochs Completed:
2 - Cumulative Sample Exposures:
30,000 - Optimizer Steps:
3,750 - Total Runtime:
~5h 42m 23s(Epoch-2 run:2h 47m 58s) - Throughput:
~1.46 - 1.49 samples/sec - Peak PyTorch VRAM:
2.63 GB - Hardware Stability: 0 OOM errors, 0 NaNs/Infs, 0 sample failures.
- Loss Progression:
- Epoch 1 Mean Loss:
~0.8655 - Epoch 2 Mean Loss:
~0.5537 - Final Step Loss:
0.4784(moving average:0.5331)
- Epoch 1 Mean Loss:
7. Benchmark Results
Evaluation was performed on a frozen benchmark suite comprising 2,300 samples per model (6,900 total evaluations across 0 failures) on standard Earth observation benchmarks: VRSBench, RSVQA-HR, and SECOND.
Multi-Model Benchmark Comparison
| Metric | Clean Zero-Shot InternVL3.5-1B | SatQuery Epoch-2 (This Checkpoint) | SatQuery Epoch-3 |
|---|---|---|---|
| VRSBench Overall | 30.28% | 43.55% | 42.69% |
| Grounding Mean IoU | 0.0434 | 0.4786 | 0.4799 |
| Grounding Median IoU | 0.0000 | 0.5262 | 0.5330 |
| Grounding IoU@0.25 | 8.00% | 75.20% | 75.60% |
| Grounding IoU@0.50 | 0.80% | 54.40% | 54.80% |
| Grounding IoU@0.75 | 0.00% | 21.60% | 20.80% |
| Grounding Coord MAE | 0.3970 | 0.0599 | 0.0598 |
| Grounding Format Validity | 58.40% | 100.00% | 100.00% |
| VRSBench Caption Score | 29.57% | 40.72% | 40.77% |
| VRSBench VQA Accuracy | 43.60% | 42.80% | 41.00% |
| SECOND Change Detection | 3.33% | 24.00% | 25.00% |
| RSVQA-HR Accuracy | 44.30% | 35.10% | 33.70% |
Latency Profile
| Metric | Clean Zero-Shot | SatQuery Epoch-2 | SatQuery Epoch-3 |
|---|---|---|---|
| Mean Latency | 1.7727s | 1.5438s | 1.4147s |
| P95 Latency | 10.4799s | 6.0242s | 5.8252s |
8. Results Interpretation & Deployment Selection
Epoch-2 was selected as the primary SatQuery deployment candidate based on multi-objective trade-offs:
- Dramatic Grounding Leap: Mean IoU increased by $+1002%$ ($0.0434 \rightarrow 0.4786$) and format validity reached $100%$.
- Optimal Balance vs Epoch-3: While Epoch-3 produced marginal additional grounding gains ($+0.0013$ Mean IoU), it showed lower aggregate VRSBench accuracy ($42.69%$ vs $43.55%$) and further degradation on VQA benchmarks ($41.00%$ vs $42.80%$).
- Diminishing Returns: Results are consistent with diminishing returns from additional epochs, where over-specialization toward bounding box generation slightly trades off general descriptive open-ended VQA reasoning.
9. Important Limitations
- Specialization Trade-off: The adapter specializes the model toward remote-sensing grounding. A modest degradation is observed on general high-resolution VQA (RSVQA-HR: $44.30% \rightarrow 35.10%$, VRSBench VQA: $43.60% \rightarrow 42.80%$).
- Approximate Visual Estimates: Bounding box coordinates represent visual estimation in image space, not sub-meter geodetic measurements. In production, coordinates should be projected through the GeoTIFF's native affine matrix.
- Resolution & Aspect Ratio: Best performance is obtained with imagery preprocessed using InternVL dynamic tiling (standard input size $448 \times 448$).
10. Exact Coordinate Convention
All spatial grounding outputs follow this convention: Where:
- $y_{\min}, x_{\min}$ is the top-left coordinate, normalized to $[0.0, 1.0]$.
- $y_{\max}, x_{\max}$ is the bottom-right coordinate, normalized to $[0.0, 1.0]$.
- Origin $(0.0, 0.0)$ is at the top-left corner of the image.
11. Quickstart & Inference
Installation
pip install torch torchvision transformers peft pillow
Python Example
import torch
from PIL import Image
from transformers import AutoModel, AutoTokenizer
from peft import PeftModel
device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.bfloat16 if device == "cuda" else torch.float32
# 1. Load base InternVL model
base_model_id = "OpenGVLab/InternVL3_5-1B-Instruct"
model = AutoModel.from_pretrained(
base_model_id,
torch_dtype=dtype,
trust_remote_code=True,
low_cpu_mem_usage=True,
).eval().to(device)
tokenizer = AutoTokenizer.from_pretrained(
base_model_id,
trust_remote_code=True,
use_fast=False,
)
# 2. Attach the SatQuery Epoch-2 LoRA adapter
adapter_id = "Praneyaarora/satquery-internvl35-1b-epoch2-lora"
if hasattr(model, "language_model"):
model.language_model = PeftModel.from_pretrained(model.language_model, adapter_id)
else:
model = PeftModel.from_pretrained(model, adapter_id)
model.eval()
# 3. Load image & run inference
# (Use the standard InternVL load_image_tensor / dynamic_preprocess pipeline)
# image = Image.open("sample_satellite.png").convert("RGB")
# pixel_values = load_image_tensor(image).to(device).to(dtype)
# prompt = "<image>\nLocate the airport runway in this satellite image."
# response = model.chat(tokenizer, pixel_values, prompt, {"max_new_tokens": 512})
# print(response)
For full preprocessing functions, see inference/inference_example.py.
12. SatQuery Integration
To use this Hugging Face adapter directly in the SatQuery environment:
In local_model_server/.env:
LOCAL_VLM_MODEL_ID=internvl-1b
LOCAL_VLM_HF_REPO=OpenGVLab/InternVL3_5-1B-Instruct
LOCAL_VLM_LORA_PATH=Praneyaarora/satquery-internvl35-1b-epoch2-lora
LOCAL_VLM_DEVICE=auto
LOCAL_VLM_PORT=8080
The server will automatically download the adapter from Hugging Face on initial startup, cache it in HF_HOME, and serve it over POST /infer.
13. License & Compliance
- Base Model License: This adapter builds upon
OpenGVLab/InternVL3_5-1B-Instruct, which is distributed under the Apache 2.0 / OpenGVLab community license. - Adapter License: Released for research, education, and geospatial application development in accordance with the base model and dataset terms of use.
14. Citation
@misc{satquery2026internvl35,
title={SatQuery: Agentic Remote-Sensing Intelligence and Spatial Grounding with Fine-Tuned Vision-Language Models},
author={Arora, Praneya and SatQuery Team},
year={2026},
publisher={Hugging Face},
howpublished={\url{https://huggingface.co/Praneyaarora/satquery-internvl35-1b-epoch2-lora}}
}
- Downloads last month
- 22
Model tree for Praneyaarora/satquery-internvl35-1b-epoch2-lora
Base model
OpenGVLab/InternVL3_5-1B-Pretrained