0Abstract
Some networks block VPN protocols; others let the connection come up and then quietly stop the traffic a few dozen seconds later. On such a network "answers the ping" and "carries traffic" are not the same thing. Colitu Adaptive Connect 2.0 is the set of mechanisms the Colitu apps use against this: server ranking that respects proximity and privacy, per-network memory kept only on the device, network hints built from anonymous counters, a connect-time check based on real traffic, server fallback, a warm spare kept ready inside the running connection, and a watcher that covers every transport. In a one-hour test with 11 faults injected at the server side, 98.8 % of traffic probes succeeded; when the server in use was cut for 101 seconds, the VPN stayed on and the user saw a stall of about 5 seconds.
1Problem and failure model
The first job of a VPN app is to reach the right server with the right connection mode. Colitu servers offer five connection modes from four protocol families: Hysteria2 (UDP/QUIC), VLESS Reality, VLESS XHTTP, Trojan and Shadowsocks 2022 (see the protocol hub).
The classic approach ranks servers by ping and picks the fastest responder. But ping is only a hint: on some networks a server answers the ping and still carries no traffic. A harder case is the silent freeze: the connection comes up, the app says "Connected", and after a while traffic stops while the TCP connection still looks open. For the user the VPN is on, but the internet is gone.
The goal of Adaptive Connect 2.0: find a path that really works when connecting, notice a path that breaks while connected, move the user to another path without turning the VPN off where possible, and do all this without collecting information about the user. The failure types below are based on behaviour seen on some networks and with some mobile operators, not on any particular country.
| Failure | What happens on the network | What the user sees |
|---|---|---|
| Protocol blocking | A protocol's connection never comes up | Cannot connect |
| Silent freezing | The connection comes up, traffic stops after N seconds, TCP stays open | "Connected", but pages don't load |
| UDP blocking | UDP-based modes (Hysteria2) don't work, TCP does | Cannot connect in the UDP mode |
| Server outage | A server can't be reached with any protocol | No connection on that server |
| Network change | Switching between Wi-Fi and mobile data, Wi-Fi dropping | Short drop; different blocking on the other network |
Three assumptions follow from this: "works" is a property of the network, not of the server; a network may freeze all TCP modes of one server together, so another TCP mode on the same server is useless as a spare; and the device's own internet can go away, in which case no mode should be blamed.
2Design
Server ranking
A location you picked is used as is; Colitu never moves a manual choice on its own. In automatic mode ("Fastest server") the order is:
| Rank | Criterion |
|---|---|
| 1 | The server that last worked on the network you are on now (remembered 24 hours) |
| 2 | Lowest measured ping; a ping older than 10 minutes or measured on another network doesn't count |
| 3 | Servers in other countries before servers in your own country, nearer countries first (also the order when there is no ping yet) |
| 4 | A server that didn't answer the ping goes to the end |
| 5 | A server that failed on this network in the last 30 minutes goes last |
"Recommended" in the server list always shows the server automatic mode would connect to. Privacy: to order the list, the panel works out your country and your internet provider's network from the IP address of that one request. Nothing is stored. When the request comes through the VPN the panel can't know them, and the app keeps the last values it learned.
Per-network memory
A "network" is the connection type (Wi-Fi, mobile data, Ethernet) combined with your internet provider's network, so home Wi-Fi and mobile data are remembered separately. The memory is kept only on the device: the last working server (24 hours), the last working mode per server (24 hours), a mode that carried no traffic (a "stall mark", 6 hours) and a server where everything failed (30 minutes). Old entries are deleted and at most 200 are kept. A mode blocked on mobile data is therefore still tried first on home Wi-Fi.
Network hints, without identifying anyone
One device's memory only knows that device's experience. Network hints share what other devices saw on the same network:
- Signed network token. The panel gives the app a token with the country and network number (ASN), signed with HMAC and valid for 48 hours. The app reports per-protocol results with it; the request's IP address is never used for observations (while connected it is a VPN server anyway).
- Anonymous hourly counters. The panel keeps only hourly counters per country, network and protocol: successes, failures and the number of distinct devices. Devices are counted with per-hour keyed pseudonyms that are deleted when the hour ends; the counts are kept for 7 days.
- Thresholds. A protocol counts as blocked on a network when at least 5 devices tried it in 24 hours with under 20 % success; with too little data, the country level is used (at least 20 devices, under 10 %). All offered protocols are never marked blocked.
- Use. The app tries blocked protocols last, unless one worked on this device on this network in the last 24 hours.
Health check and the verify path
A mode only counts as working when real traffic flows. The check asks common connectivity-test addresses (Cloudflare, Google, Microsoft) through the tunnel, never Colitu's own servers, in about 4–6 seconds per mode. Before blaming a mode, the app checks that the device is online at all. Time budget: about 8 seconds per mode, about 20 seconds per server, at most about 45 seconds for a whole connect; usually a few seconds are enough.
Verify path. Every time the VPN core starts, a private check entrance reachable only from the device itself (random port and credentials) is routed first-rule straight to the main path. The connect-time check therefore measures the main path alone: a broken main path can't hide behind the spare, and it is avoided on the next connect.
Server fallback
In automatic mode, if none of a server's modes pass the check, the app moves to the next server in the list, up to 3 servers per connect, showing "Trying another server…". The failed server stays at the back for 30 minutes on that network. With the kill switch on, Windows and Linux retry down the list instead of the same server. If every mode fails on a server you chose yourself, the error offers a one-tap "Try the fastest server".
Warm spare
While connected, the app keeps a second path ready inside the running connection. Traffic passes through a balancer: while the main path is healthy everything uses it; if it stops working, new traffic moves to the spare by itself. The VPN stays on and no "reconnecting" screen appears. Because the switch happens in the connection and not in the app screen, it also works in the background. More in the help centre: Warm spare.
| Situation | Spare |
|---|---|
| Automatic mode, a transport is proven on this network (worked in the last 24 hours, on any server) | That transport on the next server in the ranking |
| Automatic mode, nothing proven yet | The other family on the next server (UDP ↔ TCP) |
| Server chosen manually | Same server, other family; stalled modes skipped, Shadowsocks last; your location doesn't change |
| Multihop route or single-mode server | No spare |
Modes and servers that failed on this network are never the spare. With a spare attached, TCP paths use a 10-second TCP user timeout (RFC 5482) so dead connections close instead of hanging. Detection and switch usually take about 8 seconds on computers and about 10–13 seconds on phones on average (worst case about 23 seconds). Checking the main path costs about 0.3–2 MB per hour on phones and 4–7 MB per hour on computers; the spare itself only connects when it is used. The spare is also checked about once a minute through its own check entrance; a dead spare is replaced only while the tunnel is idle, never during a call. Warm spare is on by default and can be turned off in Advanced mode (Simple and Advanced mode).
Watcher and the spare-aware rule
Every transport is watched while you are connected (Adaptive Connect 1.x only watched Hysteria2): every 5 seconds for the first 90 seconds, then every 30 seconds. Three misses in a row while the device is online mean the main path is considered dead. With a spare attached, the watcher checks the main path alone and the normal path separately:
| Main path | Normal path (incl. spare) | Result |
|---|---|---|
| Working | Working | Nothing happens |
| Dead | Working: the spare carries traffic | No reconnect; the main path is marked, and the spare's transport and server lead next time |
| Dead | Dead | Reconnect |
Phones also reconnect by themselves after an unexpected drop (after about 2, 5 and 15 seconds, then the error is shown) and when the network comes back, but never after you press disconnect, sign out, or when the plan or device is paused.
Stall marks that cannot lock a network
Wrong stall marks could leave a network with nothing to try. Two rules prevent that: if the marks would leave at most one offered transport, they are ignored for that round; and marks found in a round where everything failed last only 10 minutes. The 6-hour lifetime applies only when another transport then carried traffic on the same network, that is, when the problem really was that mode.
Parallel connect
When connecting, the best two candidates are checked at the same time and the first to carry traffic wins, following the Happy Eyeballs idea of RFC 8305.
Server load
Every server reports how many people are connected to it, so a full server isn't picked and load-based distribution uses real numbers. Only complete reports are used; if a report is incomplete, the previous value is kept.
3Evaluation
Method
| Item | Detail |
|---|---|
| Device | One Android 11 phone |
| Network | One residential Wi-Fi network with active deep packet inspection (DPI) |
| Dates | 9–10 October 2026 |
| Fault injection | At the server, dropping only this client's packets: one protocol on all servers, or one whole server |
| Safety | Rules are tagged and removed by self-deleting timers; no other user is touched |
| Traffic probe | An HTTP request through the tunnel every 5–10 seconds |
Device experiments
Ranking. A server in the user's own country had the lowest ping (13–15 ms) but was correctly placed after a nearby foreign server (16–17 ms). In Adaptive Connect 1.x the list came in database order, and a server of about 222 ms on another continent was shown as "Recommended".
Silent freeze. On this network every TCP transport to one nearby server (VLESS Reality, VLESS XHTTP, Trojan) froze about 30–50 seconds after connecting, while Hysteria2 kept working. The old watcher, which only covered Hysteria2, kept showing "Connected" with no traffic. The new all-transport watcher detected the freeze and switched transport in about 1.3 seconds.
| # | Configuration | Result |
|---|---|---|
| 1 | Same server, spare Trojan, Hysteria2 cut | The spare carried traffic at +12 s; then the old watcher restarted the connection and broke it. Fixed by the spare-aware rule |
| 2 | Same server, after the fix | No TCP path survived on that server on this network; recovery came about 48 s after the block ended, because no working path existed during it |
| 3 | Automatic mode, spare on the next server as "other family" (Trojan) | Trojan froze too; led to the proven-transport rule |
| 4 | Automatic mode, proven transport: Hysteria2 on server A, spare Hysteria2 on server B, main path cut for 120 s | One failed probe at +5 s, traffic from +10 s to the end, no reconnect, the VPN stayed on. Outage ≈ 5–10 s |
One-hour fault-injection test
On 10 October 2026, a build with every 2.0 mechanism ran for 60 minutes in automatic mode on the same phone and network. An internal fault-lab tool injected 11 faults on 8 servers in random order, 2–5 minutes apart and 60–120 seconds each: either one protocol cut on every server for this client, or one server fully cut. An HTTP probe went through the tunnel every 5 seconds. 741 of 750 probes succeeded (98.8 %).
| # | Fault | Length | Longest outage | Traffic back after the fault ended |
|---|---|---|---|---|
| 1 | VLESS Reality cut on all servers | 83 s | 1 s | |
| 2 | One server in Germany fully cut | 95 s | 3 s | |
| 3 | The server in use (main path) fully cut | 101 s | 5 s | |
| 4 | Shadowsocks cut on all servers | 102 s | 4 s | |
| 5 | VLESS Reality cut on all servers | 86 s | 4 s | |
| 6 | One server in Finland fully cut | 65 s | 1 s | |
| 7 | Shadowsocks cut on all servers | 92 s | 2 s | |
| 8 | Trojan cut on all servers | 65 s | 1 s | |
| 9 | One server in Estonia fully cut | 120 s | 0 s | |
| 10 | Another server in Finland fully cut | 95 s | 1 s | |
| 11 | Hysteria2 cut on all servers | 61 s | 2 s |
- 0 s · server in use cut
- ~5 s stall
- Traffic over the warm spare
- 101 s · fault removed
The VPN stays on; no reconnect
Reading the results. Fault 3 is the key case: the server in use disappeared for 101 seconds and the user saw a stall of about 5 seconds, with the VPN on throughout. Faults on protocols and servers that were not in use show 0 seconds, as expected: a fault elsewhere doesn't disturb the session. In fault 11 every TCP protocol was already frozen by DPI on this network, so with Hysteria2 cut everywhere no working protocol existed at all; traffic returned 2 seconds after the fault ended.
4Limitations
- Narrow evaluation. One device, one network, one hour; faults were injected at the server, not by a real filter. The results can't be generalised to every network. Switching between Wi-Fi and mobile data and separate measurements on iOS, Windows and Linux devices are still to come.
- Open TCP streams. During a switch to the spare, calls continue after a short glitch and pages and messages reconnect on their own, but a large download running at that moment may restart.
- Exit rotation forces one protocol on one server and turns the warm spare off; on the measured network that meant repeated freezes.
- No spare with multihop routes or on single-mode servers.
- Stale network key. While the VPN is on, the panel can't see which network a request comes from; the app uses the last network it learned until the list is fetched again with the VPN off.
5Future work: the CSL session layer
The Colitu Session Layer (CSL) is a draft session layer planned on top of the warm spare; no date or result is promised. Today the spare moves new connections; open downloads carried over TCP can break. With CSL the app would open one session that survives when the tunnel underneath changes. In the lab prototype the longest pause was 0.08–1.75 seconds and downloads continued unbroken. A CSL session is bound to one server, so CSL handles problems within a server (network change, a blocked protocol, a short drop) and the warm spare still covers a server going down; the two work together. It would arrive as an optional "Seamless session (beta)" in Advanced mode.
6Conclusion
Real traffic, not ping, decides whether a path works, and that answer depends on the network. Ranking that respects proximity and privacy, memory kept on the device, anonymous network hints and a check that measures the main path alone find the right path; a warm spare and a watcher that knows about it turn a lost server into a pause of a few seconds.
·References
- RFC 8305, Happy Eyeballs Version 2: Better Connectivity Using Concurrency.
- RFC 9000, QUIC: A UDP-Based Multiplexed and Secure Transport.
- RFC 5482, TCP User Timeout Option.
- Hysteria2 project documentation.
- Xray-core project documentation.