Foresight: Planning Future Perception in Streaming VLMs without Retraining

Ashok Prasad Neupane1,*,†ORCID Dipan Bartaula2,* Ankit Belbase1 Saugat Adhikari1 Samip Ghimire1 Saroj Poudel1 Binod Bhattarai2,3,4 Danda Pani Paudel2,5
1Independent Researcher 2NAAMII, Nepal 3University College London, UK 4University of Aberdeen, UK 5INSAIT, Sofia University “St. Kliment Ohridski”, Bulgaria
*Equal contribution    †Corresponding author: neupane.ashok.9696@gmail.com
Paper Code arXiv

Key facts

TL;DR: Foresight makes streaming vision-language models (video LLMs) proactive without any training: a second copy of the same frozen model runs ahead of the video stream, predicts what will happen, and plans when to look, what to check and how many frames to sample.

  • Task: proactive, real-time streaming video understanding (streaming VLM / online video LLM).
  • Training-free: frozen Qwen3-VL-8B backbone; only response thresholds are calibrated.
  • Architecture: two Siamese LLMs sharing weights, encoders and one KV cache; the Ingest LLM writes, the Think LLM plans asynchronously from a snapshot.
  • OmniPro Online: 23.0 mean joint F1, best overall (MiniCPM-o 4.5: 13.5).
  • StreamingBench: 66.0 overall (+6.7 over backbone). OVO-Bench: +15.4 over backbone, +18.7 on Forward Active Responding.
  • Paper: arXiv:2610.03123 · Code: github.com/thenaivekid/foresight
Foresight teaser 23.0 joint F1
Foresight shifts streaming VLMs from reactive to proactive inference. Left: existing methods often require additional training, where the proactive and asynchronous nature are vital for human-AI interactivity. Right: Foresight performs better than training based methods while being proactive, asynchronous, and superior to real time inference on the challenging OmniPro benchmark.
23.0
mean joint F1 on OmniPro Online, best overall, training-free
+6.7
over the Qwen3-VL-8B backbone on StreamingBench
+15.4
over the backbone on OVO-Bench
0
training steps: frozen backbone

Abstract

Existing streaming vision-language models (VLMs) continuously perceive and reason over visual streams, but their computational pathways remain fixed throughout inference. Consequently, they cannot adapt computation to evolving scene dynamics, where different future events demand different levels and forms of perception. We show that streaming VLMs inherently possess the ability to anticipate the immediate future, and leverage this capability to dynamically configure future computation in a training-free manner. Realizing such anticipatory computation, however, is very challenging: future anticipation must be sufficiently reliable to guide computation, planning must run concurrently with streaming inference, and online reconfiguration must incur negligible overhead.

To address these challenges, we introduce Foresight, a dual-stream architecture comprising two Siamese LLMs with shared weights, input encoders, and KV cache. The first LLM continuously processes incoming tokens, while the second runs ahead of the stream to anticipate future context, plan future computation, and generate task responses without interrupting streaming inference. Each plan decides when to reason next, what to check then, and how densely to sample, keeping transient evidence separate from persistent control. The resulting computation plan is executed online through an efficient reconfiguration protocol with schema-guided decoding and lightweight diff-based updates, enabling dynamic adaptation with low overhead. With a frozen Qwen3-VL-8B backbone, Foresight achieves 23.0 mean joint F1 on OmniPro Online evaluation beating strongest trained baseline by 9.5%, while improving the backbone by 6.7 on StreamingBench and 15.4 on OVO-Bench, with the largest gain of 18.7 when evidence arrives later in the video stream.

Method

Foresight architecture
Foresight overview. One frozen VLM runs as two twins over a shared KV cache Ct. The Ingest LLM (perception) appends every frame admitted by the Vision Gate to Ct and never stops. When a check is due, the Think LLM (planning) reads a snapshot of Ct and writes a plan πt = (st, ut): an evidence state st and a control update ut. The Live Executor applies ut along the green control lines and emits the answer once the evidence is sufficient.

Foresight on one stream

Foresight dynamic behavior on one stream
Task: “Let me know when a player scores a goal.” Time runs left to right, and the red line marks the goal. The Ingest never pauses, and the cache grows slowly at 1 fps and faster at 8 fps. The Think LLM wakes only at scheduled checks (green bars), and the interval Δ is set by the previous plan. Early plans track only the player and ball. As a shot develops, a plan raises sampling to 8 fps and adds the goalkeeper and net; later plans compact the cache (thus KV-cache drops). Just after the goal, the evidence is sufficient and the answer is emitted; the plan then drops the goalkeeper and net, and sampling returns to 1 fps.

Results

Joint F1 (%) on the OmniPro Online subset that does not need audio. Click methods to show/hide them; hover bars for values.

MethodParamsEvent-AlertTarget-GroundState-MonitorSnap.-CountCond.-AlertCum.-CountEvent-Narr.Dedup.-CountStep-Inst.Mean
LiveStar8B9.70.80.00.014.70.01.60.06.03.6
MMDuet23B12.55.314.911.221.45.33.712.714.711.3
MiniCPM-o 4.59B50.112.69.88.023.89.03.03.81.413.5
Dispider†7B2.30.00.00.08.50.34.50.19.72.8
StreamAgent8B6.81.44.01.911.30.93.78.17.15.0
QueryStream8B8.21.43.11.017.66.67.27.18.36.7
Foresight8B44.717.916.419.133.816.311.529.218.523.0

StreamingBench accuracy (%). Click a column header to sort.

MethodFPSRTVUOmni-SourceContextualOverall
VideoLLM-Online235.9928.4526.5532.48
Flash-VStream123.2326.0024.1224.04
Dispider167.6335.6633.6153.12
TimeChat-Online‡175.3636.9635.6358.11
StreamAgent174.2836.2634.6257.02
Qwen3-VL-8B-Instruct*–74.1041.9039.8059.30
StreamBridge, LLaVA-OV168.3926.6031.7450.96
StreamBridge, LLaVA-OV + Stream-IT170.9226.3029.2151.73
StreamBridge, Qwen2-VL172.0131.3036.5755.09
StreamBridge, Qwen2-VL + Stream-IT177.0424.1032.5555.39
StreamForest‡177.1438.0536.3460.80
MiniCPM-o 4.5–78.2042.1044.5062.70
Foresightadaptive79.3048.7043.8066.00

OVO-Bench accuracy (%). Click a column header to sort.

MethodFramesReal-TimeBackwardForwardOverall
Humannative93.2092.3092.9092.80
Flash-VStream1 fps28.3727.3845.0933.61
Dispider1 fps54.5536.0634.7241.78
TimeChat-Online1 fps58.6042.0036.4045.60
StreamForest1 fps61.2052.0253.4955.57
QueryStream–61.4042.1039.0347.51
StreamAgent1 fps61.3041.7045.4049.40
Qwen3-VL-8B-Instruct*–61.6041.0037.7046.80
MiniCPM-o 4.51 fps67.3055.9055.9059.70
StreamBridge, Qwen2-VL-7B + Stream-IT1 fps71.3068.0548.3662.57
Foresightadaptive75.0255.1056.3762.16

Numbers are taken from the tables in the paper. See the paper for footnotes and evaluation details.

Ablations

Controller and input-processing ablations on OmniPro Online (%). Plan-field rows remove the named control while keeping everything else fixed. The input-processing row replaces two-frame clips with individual-frame inputs.

ComponentAblationTime F1Joint F1
Full Foresight (as reported)43.723.0
Think LLMNext Check43.018.9
Think LLMNext Question43.319.4
Think LLMNext Check + Next Question43.218.6
Vision GateFPS43.523.0
Vision EncoderKeep Classes42.722.7
Vision EncoderIgnore Classes43.522.9
KV CacheCompact Now44.323.4
Live Executorall fields (plan ignored)4.70.7
User Responseall but Answer4.00.3
Input ProcessingIndividual frames (no two-frame clips)40.119.3

Analysis

Anticipation quality against lead time
Anticipation quality against lead time.
Operating point analysis
Operating-point analysis.

FAQ: training-free proactive streaming video understanding

Foresight is a streaming video LLM / streaming VLM method for proactive, real-time video understanding. It relates to online video LLMs (VideoLLM-online, MMDuet, Dispider, StreamBridge, StreamAgent, LiveStar, StreamMind, Flash-VStream) and is evaluated on OmniPro, StreamingBench and OVO-Bench.

What is Foresight?

Foresight is a training-free, dual-stream architecture for streaming video LLMs. Two Siamese copies of one frozen VLM share weights, input encoders and KV cache: one keeps ingesting the video stream, the other runs ahead to anticipate the future and plan perception.

Does Foresight need retraining or fine-tuning?

No. The Qwen3-VL-8B backbone stays frozen; only response thresholds are calibrated on a 20% split.

How does a streaming VLM decide when to respond proactively?

In Foresight each plan decides when to reason next, what question to check then, how densely to sample frames (FPS), which objects to track or ignore, and whether there is enough evidence to answer.

How does planning avoid blocking perception?

Only the Ingest LLM writes to the shared KV cache; the Think LLM plans from a read-only snapshot, so perception and reasoning run asynchronously.

How does Foresight perform on OmniPro?

It reaches 23.0 mean joint F1 on OmniPro Online (no-audio subset), versus 13.5 for MiniCPM-o 4.5, the strongest trained baseline.

How does it do on StreamingBench and OVO-Bench?

66.0 overall on StreamingBench (+6.7 over the backbone, best in the paper's table) and +15.4 on OVO-Bench, including +18.7 on Forward Active Responding.

How is Foresight different from StreamAgent, Dispider or StreamBridge?

Dispider and StreamBridge rely on separately trained decision modules; StreamAgent plans on a fixed schedule. Foresight uses one frozen model with a shared KV cache, and its own plan sets the next check time.

Can an offline VLM become a proactive streaming assistant without training?

Yes, that is Foresight's claim: streaming VLMs already anticipate the near future, and Foresight turns that into planned, asynchronous computation without any training.

What are the limitations?

Results depend on the backbone's anticipation and temporal grounding; after a false trigger it may keep reporting the event; and it is sensitive to prompt wording.

BibTeX

@article{neupane2026foresight,
  title   = {Foresight: Planning Future Perception in Streaming VLMs without Retraining},
  author  = {Neupane, Ashok Prasad and Bartaula, Dipan and Belbase, Ankit and Adhikari, Saugat and
             Ghimire, Samip and Poudel, Saroj and Bhattarai, Binod and Paudel, Danda Pani},
  journal = {arXiv preprint arXiv:2610.03123},
  year    = {2026},
  url     = {https://arxiv.org/abs/2610.03123}
}