Skip to content

The server chosen by balancer is NOT the "best" #2030

Description

@robot-dot-win

I have a local transparent proxy with the balancer:

{
    "locals": [
        {
            "local_address": "127.0.0.1",
            "local_port": 61082,
            "protocol": "redir",
            "tcp_redir": "redirect"
        }
    ],

    "servers": [
        // The domestic server(local host):
        {
            "address": "127.0.0.1",
            "port": 51080,
            "method": "none",
            "tcp_weight": 1.0
        }

        // The overseas server:
        {
            "address": "65.x.x.x",
            "port": xxxx,
            "method": "aes-256-gcm",
            "password": "password_string",
            "tcp_weight": 0.6
        },
    ],

    "balancer": {
        "max_server_rtt": 3,
        "check_interval": 120,
        "check_best_interval": 5
    },

    "mode": "tcp_only",
    "no_delay": false,
    "fast_open": true
}

I hope most traffic will be via the domestic server, and the overseas server will act as a backup. At the beginning the chosen server is domestic, but after a while the overseas server will be chosen and from then on it seems that the balancer will always think the overseas server as the best. See log please:

Image What did it happen and how to config? Thanks.

Activity

  1. zonyitoo commented on Oct 15, 2025

    @zonyitoo
    Collaborator

    Your "domestic server" has higher latency and not stable (standard deviation is larger), which is why balancer choose another server that has lower latency and more stable.

    You could lower the weight from 0.6 to 0.01, which may help.

  2. robot-dot-win commented on Oct 15, 2025

    @robot-dot-win
    Author

    You could lower the weight from 0.6 to 0.01, which may help.

    Thanks, setting weight to 0.01 really works as expected.

    But I still wonder how it judges which is the best. In fact under most conditions the local host(domestic) will performance much better than the overseas, 'coz most destinations are domestic.

  3. database64128 commented on Oct 15, 2025

    @database64128
    Contributor

    Shameless plug: In shadowsocks-go, I implemented the following client selection policies:

    • "round-robin": Selects clients sequentially, cycling through the list.
    • "random": Selects clients randomly.
    • "availability": Selects the client with the highest availability.
    • "latency": Selects the client with the lowest average latency.
    • "min-max-latency": Selects the client with the lowest maximum latency.

    Sounds like "availability" is what you actually want. You can also customize the destination address to probe.

    https://github.com/database64128/shadowsocks-go/blob/931b4033013a2df05471d11becf2e97ce4dce57f/docs/config.json#L542-L572

  4. zonyitoo commented on Oct 16, 2025

    @zonyitoo
    Collaborator

    You could lower the weight from 0.6 to 0.01, which may help.

    Thanks, setting weight to 0.01 really works as expected.

    But I still wonder how it judges which is the best. In fact under most conditions the local host(domestic) will performance much better than the overseas, 'coz most destinations are domestic.

    As you can see in your log, your local server's latency is higher.

  5. robot-dot-win commented on Oct 17, 2025

    @robot-dot-win
    Author

    As you can see in your log, your local server's latency is higher.

    Yes, that's what i'm wondering: how does it judge the latency? The RTT of ICMP Echo Reply from the overseas server is much much bigger than from the domestic server.

  6. zonyitoo commented on Oct 17, 2025

    @zonyitoo
    Collaborator

    /// Detect TCP connectivity with Firefox's http://detectportal.firefox.com/success.txt
    async fn check_request_tcp_firefox(&self) -> io::Result<()> {
    use std::io::{Error, ErrorKind};
    const GET_BODY: &[u8] =
    b"GET /success.txt HTTP/1.1\r\nHost: detectportal.firefox.com\r\nConnection: close\r\nAccept: */*\r\n\r\n";
    let addr = Address::DomainNameAddress("detectportal.firefox.com".to_owned(), 80);
    let mut stream = ProxyClientStream::connect_with_opts(
    self.context.context(),
    self.server.server_config(),
    &addr,
    self.server.connect_opts_ref(),
    )
    .await?;
    stream.write_all(GET_BODY).await?;
    let mut reader = BufReader::new(stream);
    let mut buf = Vec::new();
    reader.read_until(b'\n', &mut buf).await?;
    let mut headers = [httparse::EMPTY_HEADER; 1];
    let mut response = httparse::Response::new(&mut headers);
    if response.parse(&buf).is_ok() && matches!(response.code, Some(200) | Some(204)) {
    return Ok(());
    }
    Err(Error::new(
    ErrorKind::InvalidData,
    format!(
    "unexpected response from http://detectportal.firefox.com/success.txt, {:?}",
    ByteStr::new(&buf)
    ),
    ))
    }

    It makes a HTTP request through shadowsocks' channel to http://detectportal.firefox.com/success.txt and checks the whole process time.

  7. robot-dot-win commented on Oct 17, 2025

    @robot-dot-win
    Author

    It makes a HTTP request through shadowsocks' channel to http://detectportal.firefox.com/success.txt and checks the whole process time.

    This doesn't seem reasonable, 'coz what users consider most is the ssservers' accessibility instead of the firefox site's accessibility! You do not need to care about the destinations' accessibility - that's what the guys who set up all the ssservers should care about.

  8. robot-dot-win commented on Oct 17, 2025

    @robot-dot-win
    Author
    • "round-robin": Selects clients sequentially, cycling through the list.
    • "random": Selects clients randomly.
    • "availability": Selects the client with the highest availability.
    • "latency": Selects the client with the lowest average latency.
    • "min-max-latency": Selects the client with the lowest maximum latency.

    Sounds like "availability" is what you actually want. You can also customize the destination address to probe.

    Strongly recommend the banlancer to support these policies!

  9. zonyitoo commented on Oct 17, 2025

    @zonyitoo
    Collaborator

    It makes a HTTP request through shadowsocks' channel to http://detectportal.firefox.com/success.txt and checks the whole process time.

    This doesn't seem reasonable, 'coz what users consider most is the ssservers' accessibility instead of the firefox site's accessibility! You do not need to care about the destinations' accessibility - that's what the guys who set up all the ssservers should care about.

    Balancer needs a common endpoint to estimate all servers' latency. Firefox's success.txt is a good choice, which is accessible from any countries by CDN.

    shadowsocks' protocol doesn't have a ping command, so there is no way to know the exact latency of the tunnel itself.

    I would accept PRs if anyone could help to implement all these policies in balancer.

  10. robot-dot-win commented on Oct 17, 2025

    @robot-dot-win
    Author

    Balancer needs a common endpoint to estimate all servers' latency. Firefox's success.txt is a good choice, which is accessible from any countries by CDN.

    Understood. But a user-defined endpoit for each ssserver seems much better, even though.

    I would accept PRs if anyone could help to implement all these policies in balancer.

    Thank you so much for your comment, and appreciate this feature very much.

  11. nouman-tariq commented on Jul 15, 2026

    @nouman-tariq

    Hi @zonyitoo you said you'd accept PRs for the balancer policies, and I'd like to help.

    Should I start with just an availability policy (opt-in, picks the server with the lowest failure rate which solves this issue and keeps the current default unchanged)? Or would you prefer the full set in one PR (round-robin, random, availability, latency, min-max-latency)?

    Let me know what works best and I'll get started.

  12. zonyitoo commented on Jul 15, 2026

    @zonyitoo
    Collaborator

    The implementation could be separated into at least these steps:

    1. Refactor to make an extensible Balancer APIs that allow adding policies easily by submitting sub-modules.
    2. Implements the current policy as the first supported Balancer.
    3. Implements each Balancer policies in separate PRs then we can have further discussion about the policy details.
  13. nouman-tariq commented on Jul 16, 2026

    @nouman-tariq

    Thanks for breaking it into steps.

    Here's my plan as I understand it:

    • Add a Balancer trait (pick the best TCP/UDP server, reset servers) that new policies can plug into.
    • Move the existing active-probing balancer behind that trait as the first implementation, with no change in behavior.
    • Keep new policies out of this first PR.

    One thing I'd like to check before I start: should the whole balancer be the pluggable part?

    I'm leaning that way because some policies (like round-robin or random) don't measure server latency at all, so they'd fit more cleanly as their own simple implementations.
    The other option is to keep one shared measuring/probing engine and only make the final "which server is best" step pluggable.

    Which fits what you had in mind?

  14. zonyitoo commented on Jul 17, 2026

    @zonyitoo
    Collaborator

    Good point. So there should be a ServerChooser or ServerManager, ServerPicker trait that allow user to pick a suitable server for a connection. Implement a BalancedServerPicker that implements ServerPicker and accept a struct impls ServerBalancer trait, then you can support different strategy/policy in different levels.

  15. nouman-tariq commented on Jul 17, 2026

    @nouman-tariq

    Hello @zonyitoo, I am going with ServerPicker/BalancedServerPicker/ServerBalancer as you specified.

    One more question

    trait ServerBalancer {
        fn best_index(&self, servers: &[ServerIdent], server_type: ServerType) -> usize;
    }
    

    Today's scoring only reads a plain atomic value, so sync works fine. But a future availability policy will need richer stats that sit behind an async Mutex. Two options:

    1. Make best_index async, so future balancers can await the mutex directly.
    2. Or keep it sync, and have the BalancedServerPicker snapshot the stats into plain data before calling it.

    Which would you prefer?

  16. UbiVPN commented on Jul 18, 2026

    @UbiVPN

    Ran into this too. The trick was finding a transport that looks like normal HTTPS. I use a vpn that has this built in, but people setting up manually could achieve the same with the right config.

  17. zonyitoo commented on Jul 18, 2026

    @zonyitoo
    Collaborator

    The ServerPicker trait should also manage all the server instances:

    trait ServerPicker {
        // Initialize / Reinitialize all server instances
        async fn set_servers(&self, servers: impl IntoIter<Server>, server_type: ServerType);
    
        // Called by the server Client when it needs a server instance to connect
        async fn pick_server(&self, server_type: ServerType) -> Result<Arc<Server>, PickError>;
    }

    Simple policies, like round-robin, impls ServerPicker would be quite straight forward.

    And then make a BalancedServerPicker that impls ServerPicker:

    struct ScoreBalancedServerPicker<B>
        where B: ServerScoreBalancer
    {
        servers: Vec<Server>,
        server_balancer: B,
    }
    
    impl<B> ServerPicker for ScoreBalancedServerPicker<B> { ... }

    The ServerScoreBalancer provide methods to get Server's score:

    trait ServerScoreBalancer {
        type Error;
    
        // Called by the ScoreBalancedServerPicker when it needs a server's score
        // The score is [0, 1], the higher the better
        async fn get_server_score(&self, server: &Server, server_type: ServerType) -> Result<f64, Error>;
    }

    The names are not finalized.

    What do you think?

  18. UbiVPN commented on Jul 24, 2026

    @UbiVPN

    Had the same issue in Russia. What solved it was tunneling through a protocol that randomises packet sizes. Took some trial and error. There are services that do this out of the box now (ubivpn is one I know).

  19. laikitleung commented on Aug 2, 2026

    @laikitleung

    A balancer is just a balancer,why does it need to be the best?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions