很多人把 worker_connections 从 768 改成 65535 后就以为并发上去了,压测发现一点没变——因为真正的天花板往往不在 Nginx,而在系统参数和后端。这篇把「Nginx 配置 + 系统内核 + 压测定位」三层串起来讲。
一、先算清楚账:并发到底卡在哪
理论最大连接数 = worker_processes × worker_connections
反向代理场景下每个客户端连接还要占一个到后端的连接,所以实际承载的客户端并发要除以 2。4 核机器配 10240,静态资源场景约 4 万连接,反向代理约 2 万。
三个常见天花板(按出现频率排序):
| 天花板 | 现象 | 查法 |
|---|---|---|
| 系统 nofile | error.log 里 Too many open files / worker_connections are not enough |
cat /proc/<nginx-pid>/limits |
| somaxconn 队列 | 压测出现连接超时、ss -s 里 synrecv 堆积 |
ss -lnt 看 Recv-Q,netstat -s | grep listen |
| 后端扛不住 | Nginx 正常但 502/504 增多 | 看 upstream_response_time |
二、Nginx 层配置
user www-data;
worker_processes auto;
worker_rlimit_nofile 65535; # 必须!否则 worker_connections 配再大也被 ulimit 卡住
events {
use epoll;
worker_connections 10240;
multi_accept on;
accept_mutex off; # 高并发下关掉更均匀(Linux 4.5+ 有 EPOLLEXCLUSIVE)
}
http {
sendfile on;
tcp_nopush on;
tcp_nodelay on;
keepalive_timeout 30; # 长连接别太长,否则空闲连接占着 worker
keepalive_requests 10000;
server_tokens off;
client_max_body_size 20m;
client_body_timeout 10s;
client_header_timeout 10s;
send_timeout 30s;
}
worker_rlimit_nofile 是最容易被漏掉的一行——它直接改 worker 进程的文件句柄上限,不写这行,ulimit 还是 1024,后面全白配。
三、系统层配套(不改等于没调)
# /etc/security/limits.conf
www-data soft nofile 65535
www-data hard nofile 65535
# systemd(Ubuntu/CentOS 7+ 必须走这条)
mkdir -p /etc/systemd/system/nginx.service.d
cat > /etc/systemd/system/nginx.service.d/limits.conf <<'EOF'
[Service]
LimitNOFILE=65535
EOF
systemctl daemon-reload && systemctl restart nginx
# /etc/sysctl.d/99-nginx.conf
net.core.somaxconn = 65535
net.ipv4.tcp_max_syn_backlog = 65535
net.ipv4.ip_local_port_range = 1024 65535
net.ipv4.tcp_tw_reuse = 1
net.ipv4.tcp_fin_timeout = 15
net.core.netdev_max_backlog = 65535
并且 listen 要带 backlog:
listen 80 backlog=65535;
listen 443 ssl http2 backlog=65535;
四、反向代理场景的关键优化
upstream backend {
server 127.0.0.1:8080 max_fails=3 fail_timeout=10s;
keepalive 64; # 到后端的空闲连接池
}
server {
location /api/ {
proxy_pass http://backend;
proxy_http_version 1.1; # upstream keepalive 的前提
proxy_set_header Connection "";
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_connect_timeout 3s;
proxy_read_timeout 30s;
proxy_send_timeout 30s;
proxy_next_upstream error timeout http_502;
proxy_next_upstream_tries 2;
}
}
keepalive 64 的价值:没它时每个请求都要和后端重新 TCP 握手,QPS 上不去且后端 TIME_WAIT 爆炸。经验值(估算):后端实例数 × 并发峰值 / 实例数,先给 64,压测后调。
五、加日志字段,才能定位瓶颈
默认日志看不出慢在哪,加两个字段:
log_format main '$remote_addr $request_method $uri $status '
'rt=$request_time urt=$upstream_response_time '
'uct="$upstream_connect_time" $body_bytes_sent';
access_log /var/log/nginx/access.log main;
判据很简单:
rt大、urt小 → 慢在 Nginx 自己(多半是客户端网络或大文件发送)rt大、urt也大 → 慢在后端,去查应用和数据库urt是-→ 请求根本没到后端(502 场景)
日志有了,直接贴进 日志分析工具,5xx、超时、连接拒绝会按类型归类出来。
六、压测怎么测
# 安装 wrk(比 ab 更能压出真实并发)
apt install -y wrk
# 12 线程 / 400 并发 / 压 30 秒
wrk -t12 -c400 -d30s --latency http://127.0.0.1/
看三个数:Requests/sec(QPS)、Latency 分布、Non-2xx 数量。压测时同时在服务器上跑:
watch -n1 'ss -s; cat /proc/net/sockstat | head -3'
top -H -p $(pgrep -f "nginx: worker" | head -1)
如果 QPS 上不去但 CPU 空闲 → 多半卡在后端或数据库连接;如果 worker 进程 CPU 打满 → 该加 worker 进程或上多机;如果 ss -s 里 TIME_WAIT 几万 → 检查 tcp_tw_reuse 和 upstream keepalive。
七、一份可直接抄的最小配置结论
| 场景 | worker_connections | 系统 nofile | somaxconn | upstream keepalive |
|---|---|---|---|---|
| 静态资源 / CDN 回源 | 10240 | 65535 | 65535 | — |
| 反向代理(4 核 8G) | 10240 | 65535 | 65535 | 64 |
| 小内存(1C2G) | 4096 | 20480 | 8192 | 32 |
数值均为经验估算起点,以压测结果为准;改一项压一次,别一次改一堆,否则不知道哪个起了作用。
相关工具:HTTP 健康检查 · HTTP 状态码查询 · 日志分析