FlyVis OPTP experimental demo · 10 seconds · No audio
On the left are a fly model and input images. The middle compares page classifications from the existing and new models. The right visualizes features and projections along the horizontal and vertical axes. As the input changes, the two models’ predictions can be compared. Browser timings shown in the recording may differ from the separate internal measurements discussed below.
Can a smaller model see better?
Recognizing pages is a basic part of a book-scanning app. The vFlat development team experimented with making the OPTP page recognition model lighter. An unexpected source of inspiration was research into fruit-fly vision.
This article describes a research and development experiment shared with the team on September 15, 2026. It is not an announcement of a released vFlat feature or performance verified across phones.
Borrowing an idea, not copying a nervous system
The developer referred to FlyVis, but did not copy an entire fruit-fly neural circuit. Instead, the experiment adopted an idea of compressing information separately along the horizontal and vertical directions into one-dimensional representations, adapted for page recognition.
The existing model extracted features while progressively reducing a two-dimensional image. The new model first extracts simple two-dimensional features quickly, then compresses them along the two axes for processing. The aim is to reduce the work of continuing to compute over the full spatial representation.
Half the parameters, higher internal test accuracy
In the developer’s internal test, accuracy increased from 96.5% to 98.6%, a gain of 2.1 percentage points. The parameter count was roughly halved, and the computational workload fell to about one third.
The report also recorded latency below 1 ms, with a minimum of 0.44 ms. These results belong to that experimental setup. The measurement device, runtime, repetition count and test-set composition are not disclosed here, so these numbers should not be generalized to speed on individual phones or accuracy in everyday use.
The next target: a mobile CPU
The developer’s next goal is to run the model directly on a mobile CPU, avoiding the step of moving data to a GPU. Reducing the model’s computational workload opens up room to explore this execution path.
Sufficiently fast mobile CPU performance remains a hypothesis to test. Latency on real devices and accuracy across different shooting conditions still need validation.
A bigger model is not the only direction
What makes this experiment interesting is that the internal test result improved without making the model larger. A developer’s effort to simplify an idea from another field and adapt it to a specific problem has shown a possible route to lighter page recognition.
Article: the vFlat development team. Figures and implementation details are based on the developer’s internal experimental report shared on September 15, 2026.