haproxy

mirror of https://git.haproxy.org/git/haproxy.git/ synced 2025-08-11 09:37:20 +02:00

Author	SHA1	Message	Date
Fr�d�ric L�caille	055e82657e	BUG/MINOR: quic: Do not ignore coalesced packets in qc_prep_fast_retrans() This function is called only when probing only one packet number space (Handshake) or two times the same one (Application). So, there is no risk to prepare two times the same frame when uneeded because we wanted to probe two packet number spaces. The condition "ignore the packets which has been coalesced to another one" is not necessary. More importantly the bug is when we want to prepare a Application packet which has been coalesced to an Handshake packet. This is always the case when the first Application packet is sent. It is always coalesced to an Handshake packet with an ACK frame. So, when lost, this first application packet was never resent. It contains the HANDSHAKE_DONE frame to confirm the completion of the handshake to the client. Must be backported to 2.6 and 2.7.	2023-02-03 17:55:55 +01:00
Fr�d�ric L�caille	6dead91b8a	MINOR: quic: Add a trace about variable states in qc_prep_fast_retrans() This has already been very useful to diagnose retransmission issues. Must be backported to 2.6 and 2.7.	2023-02-03 17:55:55 +01:00
Fr�d�ric L�caille	b75eecc874	BUG/MINOR: quic: Too big PTO during handshakes During the handshake and when the handshake has not been confirmed the acknowledgement delays reported by the peer may be larger than max_ack_delay. max_ack_delay SHOULD be ignored before the handshake is completed when computing the PTO. But the current code considered the wrong condition "before the hanshake is completed". Replace the enum value QUIC_HS_ST_COMPLETED by QUIC_HS_ST_CONFIRMED to fix this issue. In quic_loss.c, the parameter passed to quic_pto_pktns() is renamed to avoid any possible confusion. Must be backported to 2.7 and 2.6.	2023-02-03 17:55:55 +01:00
Fr�d�ric L�caille	dd419461ef	BUG/MINOR: quic: Possible stream truncations under heavy loss This may happen during retransmission of frames which can be splitted (CRYPTO, or STREAM frames). One may have to split a frame to be retransmitted due to the QUIC protocol properties (packet size limitation and packet field encoding sizes). The remaining part of a frame which cannot be retransmitted must be detached from the original frame it is copied from. If not, when the really sent part will be acknowledged the remaining part will be acknowledged too but not sent! Must be backported to 2.7 and 2.6.	2023-02-03 17:55:55 +01:00
Fr�d�ric L�caille	9969adbcdc	MINOR: stats: add by HTTP version cumulated number of sessions and requests Add cum_sess_ver[] new array of counters to count the number of cumulated HTTP sessions by version (h1, h2 or h3). Implement proxy_inc_fe_cum_sess_ver_ctr() to increment these counter. This function is called each a HTTP mux is correctly initialized. The QUIC must before verify the application operations for the mux is for h3 before calling proxy_inc_fe_cum_sess_ver_ctr(). ST_F_SESS_OTHER stat field for the cumulated of sessions others than HTTP sessions is deduced from ->cum_sess_ver counter (for all the session, not only HTTP sessions) from which the HTTP sessions counters are substracted. Add cum_req[] new array of counters to count the number of cumulated HTTP requests by version and others than HTTP requests. This new member replace ->cum_req. Modify proxy_inc_fe_req_ctr() which increments these counters to pass an HTTP version, 0 special values meaning "other than an HTTP request". This is the case for instance for syslog.c from which proxy_inc_fe_req_ctr() is called with 0 as version parameter. ST_F_REQ_TOT stat field compputing for the cumulated number of requests is modified to count the sum of all the cum_req[] counters. As this patch is useful for QUIC, it must be backported to 2.7.	2023-02-03 17:55:49 +01:00
Willy Tarreau	23aa79d9a9	OPTIM: htx: inline the most common memcpy(8) On high traffic benchmarks, it's visible the the CPU is dominated by calls to memcpy(), and many of those come from htx functions. It was measured that 63% of those coming from htx are made on 8-byte blocks which really are not worth a call to the function since a single read-write cycle does it fine. This commit adds an inline htx_memcpy() function that explicitly checks for this length and just copies the data without that call. It's even likely that it could be detected on const sizes, though that was not done. This is already effective in reducing the number of calls to memcpy().	2023-02-03 13:39:18 +01:00
Amaury Denoyelle	24d5b72ca9	MINOR: quic: add config for retransmit limit Define a new configuration option "tune.quic.max-frame-loss". This is used to specify the limit for which a single frame instance can be detected as lost. If exceeded, the connection is closed. This should be backported up to 2.7.	2023-02-03 11:56:46 +01:00
Amaury Denoyelle	e4abb1f2da	MEDIUM: quic: implement a retransmit limit per frame Add a <loss_count> new field in quic_frame structure. This field is set to 0 and incremented each time a sent packet is declared lost. If <loss_count> reached a hard-coded limit, the connection is deemed as failing and is closed immediately with a CONNECTION_CLOSE using INTERNAL_ERROR. By default, limit is set to 10. This should ensure that overall memory usage is limited if a peer behaves incorrectly. This should be backported up to 2.7.	2023-02-03 11:56:42 +01:00
Amaury Denoyelle	57b3eaa793	MINOR: quic: refactor frame deallocation Define a new function qc_frm_free() to handle frame deallocation. New BUG_ON() statements ensure that the deallocated frame is not referenced by other frame. To support this, all LIST_DELETE() have been replaced by LIST_DEL_INIT(). This should enforce that frame deallocation is robust. As a complement, qc_frm_unref() has been moved into quic_frame module. It is justified as this is a utility function related to frame deallocation. It allows to use it in quic_pktns_tx_pkts_release() before calling qc_frm_free(). This should be backported up to 2.7.	2023-02-03 11:55:41 +01:00
Amaury Denoyelle	40c24f1a10	MINOR: quic: define new functions for frame alloc Define two utility functions for quic_frame allocation : * qc_frm_alloc() is used to allocate a new frame * qc_frm_dup() is used to allocate a new frame by duplicating an existing one Theses functions are useful to centralize quic_frame initialization. Note that pool_zalloc() is replaced by a proper pool_alloc() + explicit initialization code. This commit will simplify implementation of the per frame retransmission limitation. Indeed, a new counter will be added in quic_frame structure which must be initialized to 0. This should be backported up to 2.7.	2023-02-03 10:44:26 +01:00
Amaury Denoyelle	1dac018d9f	MINOR: quic: ensure offset is properly set for STREAM frames Care must be taken when reading/writing offset for STREAM frames. A special OFF bit is set in the frame type to indicate that the field is present. If not set, it is assumed that offset is 0. To represent this, offset field of quic_stream structure must always be initialized with a valid value in regards with its frame type OFF bit. The previous code has no bug in part because pool_zalloc() is used to allocate quic_frame instances. To be able to use pool_alloc(), offset is always explicitely set to 0. If a non-null value is used, OFF bit is set at the same occasion. A new BUG_ON() statement is added on frame builder to ensure that the caller has set OFF bit if offset is non null. This should be backported up to 2.7.	2023-02-03 09:46:55 +01:00
Amaury Denoyelle	2216b0866e	MINOR: quic: remove fin from quic_stream frame type A dedicated <fin> field was used in quic_stream structure. However, this info is already encoded in the frame type field as specified by QUIC protocol. In fact, only code for packet reception used the <fin> field. On the sending side, we only checked for the FIN bit. To align both sides, remove the <fin> field and only used the FIN bit. This should be backported up to 2.7.	2023-02-03 09:46:55 +01:00
Aurelien DARRAGON	5e7ecbec99	BUG/MINOR: stats: use proper buffer size for http dump In an attempt to fix GH #1873, ("BUG/MEDIUM: stats: Rely on a local trash buffer to dump the stats") explicitly reduced output buffer size to leave enough space for htx overhead under http context. Github user debtsandbooze, who first reported the issue, came back to us and said he was still able to make the http dump "hang" with the new fix. After some tests, it became clear that htx_add_data_atonce() could fail from time to time in stats_putchk(), even if htx was completely empty: In http context, buffer size is maxed out at channel_htx_recv_limit(). Unfortunately, channel_htx_recv_limit() is not what we're looking for here because limit() doesn't compute the proper htx overhead. Using buf_room_for_htx_data() instead of channel_htx_recv_limit() to compute max "usable" data space seems to be the last piece of work required for the previous fix to work properly. This should be backported everywhere the aforementioned commit is.	2023-02-02 17:10:11 +01:00
Aurelien DARRAGON	739281b3d6	BUG/MEDIUM: thread: consider secondary threads as idle+harmless during boot idle and harmless bits in the tgroup_ctx structure were not explicitly set during boot. \| struct tgroup_ctx ha_tgroup_ctx[MAX_TGROUPS] = { }; As the structure is first statically initialized, .threads_harmless and .threads_idle are automatically zero- initialized by the compiler. Unfortulately, this means that such threads are not considered idle nor harmless by thread_isolate(_full)() functions until they enter the polling loop (thread_harmless_now() and thread_idle_now() are respectively called before entering the polling loop) Because of this, any attempt to call thread_isolate() or thread_isolate_full() during a startup phase with nbthreads >= 2 will cause thread_isolate to loop until every secondary threads make it through their first polling loop. If the startup phase is aborted during boot (ie: "-c" option to check the configuration), secondary threads may be initialized but will never be started (ie: they won't enter the polling loop), thus thread_isolate() could would loop forever in such cases. We can easily reveal the bug with this patch reproducer: \| diff --git a/src/haproxy.c b/src/haproxy.c \| index e91691658..0b733f6ee 100644 \| --- a/src/haproxy.c \| +++ b/src/haproxy.c \| @@ -2317,6 +2317,10 @@ static void init(int argc, char *argv) \| if (pr \|\| px) { \| / At least one peer or one listener has been found */ \| qfprintf(stdout, "Configuration file is valid\n"); \| + printf("haproxy will loop...\n"); \| + thread_isolate(); \| + printf("we will never reach this\n"); \| + thread_release(); \| deinit_and_exit(0); \| } \| qfprintf(stdout, "Configuration file has no error but will not start (no listener) => exit(2).\n"); Now we start haproxy with a valid config: $> haproxy -c -f valid.conf Configuration file is valid haproxy will loop... ^C ------------------------------------------------------------------------------ This did not cause any issue so far because no early deinit paths require full thread isolation. But this may change when new features or requirements are introduced, so we should fix this before it becomes a real issue. To fix this, we explicitly assign .threads_harmless and .threads_idle to .threads_enabled value in thread_map_to_groups() function during boot. This is the proper place to do this since as long as .threads_enabled is not explicitly set, its default value is also 0 (zero-initialized by the compiler) code snippet from thread_isolate() function: ulong te = _HA_ATOMIC_LOAD(&ha_tgroup_info[tgrp].threads_enabled); ulong th = _HA_ATOMIC_LOAD(&ha_tgroup_ctx[tgrp].threads_harmless); if ((th & te) == te) break; Thus thread_isolate(_full()) won't be looping forever in thread_isolate() even if it were to be used before thread_map_to_groups() is executed. No backport needed unless this is a requirement.	2023-02-02 08:21:15 +01:00
Amaury Denoyelle	78adb4b451	BUG/MINOR: h3: fix crash due to h3 traces This commit is identical to the preceeding patch. However, these traces are from another patch with a different backport scope : `56a86ddfb9` MINOR: h3: add missing traces on closure This must be backported up to 2.7 where above patch is scheduled.	2023-01-31 16:09:47 +01:00
Amaury Denoyelle	e31867b7fa	BUG/MINOR: h3: fix crash due to h3 traces First H3 traces argument must be a connection instance or a NULL. Some new traces were added recently with a qcc instance which caused a crash when traces are activated. This trace was added by the following patch : `87f8766d3f` BUG/MEDIUM: h3: handle STOP_SENDING on control stream This must be backported up to 2.6 along with the above patch.	2023-01-31 16:08:33 +01:00
William Lallemand	222e5a260b	BUG/MEDIUM: ssl: wrong eviction from the session cache tree When using WolfSSL, there are some cases were the SSL_CTX_sess_new_cb is called with an existing session ID. These cases are not met with OpenSSL. When the ID is found in the session tree during the insertion, the shared_block len is not set to 0 and is not used. However if later the block is reused, since the len is not set to 0, the release callback will be called an ebmb_delete will be tried on the block, even if it's not in the tree, provoking a crash. The code was buggy from the beginning, but the case never happen with openssl which changes the ID. Must be backported in every maintained branches.	2023-01-31 14:34:40 +01:00
Amaury Denoyelle	56a86ddfb9	MINOR: h3: add missing traces on closure Add traces for function h3_shutdown() / h3_send_goaway(). This should help to debug problems related to connection closure. This should be backported up to 2.7.	2023-01-30 16:16:46 +01:00
Amaury Denoyelle	e269aeb46b	BUG/MINOR: h3: reject RESET_STREAM received for control stream This commit is similar to the previous one. It reports an error if a RESET_STREAM is received for the remote control stream. This will generate a CONNECTION_CLOSE with H3_CLOSED_CRITICAL_STREAM error. Note that contrary to the previous bug related to STOP_SENDING, this bug was not encountered in real environment. As such, it is labelled as MINOR. However, it could triggered the same crash as the previous patch. This should be backported up to 2.6.	2023-01-30 16:16:46 +01:00
Amaury Denoyelle	87f8766d3f	BUG/MEDIUM: h3: handle STOP_SENDING on control stream Before this patch, STOP_SENDING reception was considered valid even on H3 control stream. This causes the emission in return of RESET_STREAM and eventually the closure and freeing of the QCS instance. This then causes a crash during connection closure as a GOAWAY frame is emitted on the control stream which is now released. To fix this crash, STOP_SENDING on the control stream is now properly rejected as specified by RFC 9114. The new app_ops close callback is used which in turn will generate a CONNECTION_CLOSE with error H3_CLOSED_CRITICAL_STREAM. This bug was detected in github issue #2006. Note that however it is triggered by an incorrect client behavior. It may be useful to determine which client behaves like this. If this case is too frequent, STOP_SENDING should probably be silently ignored. To reproduce this issue, quiche was patched to emit a STOP_SENDING on its send() function in quiche/src/lib.rs: pub fn send(&mut self, out: &mut [u8]) -> Result<(usize, SendInfo)> { - self.send_on_path(out, None, None) + let ret = self.send_on_path(out, None, None); + self.streams.mark_stopped(3, true, 0); + ret } This must be backported up to 2.6 along with the preceeding commit : MINOR: mux-quic/h3: define close callback	2023-01-30 16:12:23 +01:00
Amaury Denoyelle	1e340ba6bc	MINOR: mux-quic/h3: define stream close callback Define a new qcc_app_ops callback named close(). This will be used to notify app-layer about the closure of a stream by the remote peer. Its main usage is to ensure that the closure is allowed by the application protocol specification. For the moment, close is not implemented by H3 layer. However, this function will be mandatory to properly reject a STOP_SENDING on the control stream and preventing a later crash. As such, this commit must be backported with the next one on 2.6. This is related to github issue #2006.	2023-01-30 15:56:25 +01:00
Amaury Denoyelle	4be5435014	OPTIM: h3: skip buf realign if no trailer to encode h3_resp_trailers_send() may be called due to an HTX EOT block present without preceeding HTX TRAILER block. In this case, no HEADERS frame will be generated by H3 layer and MUX will emit an empty STREAM frame with FIN set. However, before skipping these, some operations are conducted on qcs buffer to realign it and try to encode the QPACK field section line in a buffer copy. These operation are thus unneeded if no trailer is generated. Even worse, the function will fail if there is not enough space in the buffer for the superfluous QPACK section line. To improve this situation, this patch adds an early goto statement to skip most operations in h3_resp_trailers_send() if no HTX trailer block is found. This patch is related to github issue #2006. This should be backported up to 2.7.	2023-01-30 15:39:41 +01:00
Amaury Denoyelle	224ba5cffe	BUG/MEDIUM: h3: do not crash if no buf space for trailers Replace ABORT_NOW() by proper error management in h3_resp_trailers_send() for QPACK encoding operation. If a QPACK encoding operation fails, it means there is not enough space in qcs buffer. In this case, flag qcs instance with QC_SF_BLK_MROOM and return an error. MUX is responsible to remove this flag once buffer space is available. This should fix the crash reported by gabrieltz on github issue #2006. This must be backported up to 2.7.	2023-01-30 15:38:22 +01:00
Aurelien DARRAGON	8436c910f5	BUG/MINOR: http_ext/7239: ipv6 dumping relies on out of scope variables In http_build_7239_header_nodename(), ip6 address dumping is performed at a single place to prevent code duplication: A goto statement combined with a local pointer variable (ip6_addr) were used to perform ipv6 dump from different calling places inside the function. However, when the goto was performed (ie: sample expression handling), ip6_addr pointer was assigned to limited scope variable's address that is not valid within the dumping code. Because of this, we have an undefined behavior that could result in a bug or a crash depending on the platform that is running haproxy. This was found by Coverity (GH #2018) To fix this, we add a simple ip6 printing helper that takes the ip6_addr pointer as an argument. This prevents any scope related bug as the function is executed under the proper context. if/else guards inside the function were reviewed to make sure that the goto removal won't affect existing behavior. ---------- No backport needed, except if the commit ("MINOR: proxy/http_ext: introduce proxy forwarded option") is backported. Given that this commit needs to be backported with "MINOR: proxy/http_ext: introduce proxy forwarded option", We're using it as a reminder for another bug that was introduced with "MINOR: proxy/http_ext: introduce proxy forwarded option" but has been silently fixed since with "MEDIUM: proxy/http_ext: implement dynamic http_ext". If "MINOR: proxy/http_ext: introduce proxy forwarded option" needs to be backported without "MEDIUM: proxy/http_ext: implement dynamic http_ext", you should manually apply the following patch on top of it: \| diff --git a/src/http_ext.c b/src/http_ext.c \| index fcb5a07bc..3921357a3 100644 \| --- a/src/http_ext.c \| +++ b/src/http_ext.c \| @@ -609,7 +609,7 @@ static inline void http_build_7239_header_node(struct buffer out, \| if (forby->np_mode) \| chunk_appendf(out, "\""); \| offset_save = out->data; \| - http_build_7239_header_node(out, s, curproxy, addr, &curproxy->http.fwd.p_by); \| + http_build_7239_header_nodename(out, s, curproxy, addr, forby); \| if (offset_save == out->data) { \| / could not build nodename, either because some \| * data is not available or user is providing bad input \| @@ -619,7 +619,7 @@ static inline void http_build_7239_header_node(struct buffer out, \| if (forby->np_mode) { \| chunk_appendf(out, ":"); \| offset_save = out->data; \| - http_build_7239_header_nodeport(out, s, curproxy, addr, &curproxy->http.fwd.p_by); \| + http_build_7239_header_nodeport(out, s, curproxy, addr, forby); \| if (offset_save == out->data) { \| / could not build nodeport, either because some data is \| * not available or user is providing bad input (If you don't, forwarded option won't work properly and will crash haproxy (stack overflow) when building 'for' or 'by' parameter)	2023-01-30 15:14:08 +01:00
Christopher Faulet	c254516c53	BUG/MINOR: mux-h2: Fix possible null pointer deref on h2c in _h2_trace_header() As reported by Coverity, this function may be called with no h2c. Thus, the pointer must always be checked before any access. One test was missing in TRACE_PRINTF_LOC(). This patch should fix the issue #2015. No backport needed, except if the commit `11e8a8c2a` ("MEDIUM: mux-h2/trace: add tracing support for headers") is backported.	2023-01-30 08:26:12 +01:00
Aurelien DARRAGON	d49a580fda	BUG/MINOR: fcgi-app: prevent 'use-fcgi-app' in default section Despite the doc saying that 'use-fcgi-app' keyword may only be used in backend or listen section, we forgot to prevent its usage in default section. This is wrong because fcgi relies on a filter, and filters cannot be defined in a default section. Making sure such usage reports an error to the user and complies with the doc. This could be backported up to 2.2.	2023-01-27 15:18:59 +01:00
Aurelien DARRAGON	c07001bb56	MINOR: cfgparse/http_ext: move post-parsing http_ext steps to http_ext This is a simple refactor to remove specific http_ext post-parsing treatment from cfgparse. Related work is now performed internally through check_http_ext_postconf() function that is registered via REGISTER_POST_PROXY_CHECK() in http_ext.c.	2023-01-27 15:18:59 +01:00
Aurelien DARRAGON	b2e2ec51b3	MEDIUM: proxy/http_ext: implement dynamic http_ext proxy http-only options implemented in http_ext were statically stored within proxy struct. We're making some changes so that http_ext are now stored in a dynamically allocated structs. http_ext related structs are only allocated when needed to save some space whenever possible, and they are automatically freed upon proxy deletion. Related PX_O_HTTP{7239,XFF,XOT) option flags were removed because we're now considering an http_ext option as 'active' if it is allocated (ptr is not NULL) A few checks (and BUG_ON) were added to make these changes safe because it adds some (acceptable) complexity to the previous design. Also, proxy.http was renamed to proxy.http_ext to make things more explicit.	2023-01-27 15:18:59 +01:00
Aurelien DARRAGON	d745a3f117	MINOR: http_ext/7239: warn the user when fetch is not available http option forwarded (rfc7239) supports sample expressions when configuring 'host', 'for' and 'by' parameters. However, since we are in a http-backend-only context, right after http header is processed, we have a limited resolution scope for the sample expression provided by the user. To prevent any confusion, a warning is emitted when parsing the option if the user relies on a sample expression (more precisely a fetch) which would yield unexpected results at runtime when processing the option.	2023-01-27 15:18:59 +01:00
Aurelien DARRAGON	9ded834adc	OPTIM: http_ext/7239: introduce c_mode to save some space forwarded header option (rfc7239) deals with sample expressions in two steps: first a sample expression string is extracted from the config file and later in startup sequence this string is converted into the resulting sample_expr. We need to perform these two steps because we cannot compile the expr too early in the parsing sequence. (or we would miss some context) Because of this, we have two dinstinct structure members (expr and expr_s) for each 7239 field supporting sample expressions. This is not cool, because we're bloating the http forwarded config structure, and thus, bloating proxy config structure. To address this, we now merge both expr and expr_s members inside a single union to regain some space. This forces us to perform some additional logic to make sure to use the proper structure member at different parsing steps. Thanks to this, we're also able to free/release related config hints and sample expression strings as soon as the sample expression compilation is finished.	2023-01-27 15:18:59 +01:00
Aurelien DARRAGON	9a273b4069	MINOR: http_ext: add rfc7239_n2np converter Adding new http converter: rfc7239_n2np. Takes a string representing 7239 forwarded header node (extracted from either 'for' or 'by' 7239 header fields) as input and translates it to either unsigned integer or ('_' prefixed obfuscated identifier), according to 7239RFC. Example: # extract 'by' field from forwarded header, extract node port from # resulting node identifier and store the result in req.fnp http-request set-var(req.fnp) req.hdr(forwarded),rfc7239_field(by),rfc7239_n2np #input: "by=\"127.0.0.1:9999\"" # output: 9999 #input: "by=\"_name:_port\"" # output: "_port" Depends on: - "MINOR: http_ext: introduce http ext converters"	2023-01-27 15:18:59 +01:00
Aurelien DARRAGON	07d6753c89	MINOR: http_ext: add rfc7239_n2nn converter Adding new http converter: rfc7239_n2nn. Takes a string representing 7239 forwarded header node (extracted from either 'for' or 'by' 7239 header fields) as input and translates it to either ipv4 address, ipv6 address or str ('_' prefixed if obfuscated or "unknown" if unknown), according to 7239RFC. Example: # extract 'for' field from forwarded header, extract nodename from # resulting node identifier and store the result in req.fnn http-request set-var(req.fnn) req.hdr(forwarded),rfc7239_field(for),rfc7239_n2nn #input: "for=\"127.0.0.1:9999\"" # output: 127.0.0.1 #input: "for=\"_name:_port\"" # output: "_name" Depends on: - "MINOR: http_ext: introduce http ext converters"	2023-01-27 15:18:59 +01:00
Aurelien DARRAGON	6fb58b8c9d	MINOR: http_ext: add rfc7239_field converter Adding new http converter: rfc7239_field. Takes a string representing 7239 forwarded header single value as input and extracts a single field/parameter from the header according to user selection. Example: # extract host field from forwarded header and store it in req.fhost var http-request set-var(req.fhost) req.hdr(forwarded),rfc7239_field(host) #input: "proto=https;host=\"haproxy.org:80\"" # output: "haproxy.org:80" # extract for field from forwarded header and store it in req.ffor var http-request set-var(req.ffor) req.hdr(forwarded),rfc7239_field(for) #input: "proto=https;host=\"haproxy.org:80\";for=\"127.0.0.1:9999\"" # output: "127.0.0.1:9999" Depends on: - "MINOR: http_ext: introduce http ext converters"	2023-01-27 15:18:59 +01:00
Aurelien DARRAGON	5c6f86f465	MINOR: http_ext: add rfc7239_is_valid converter Adding new http converter: rfc7239_is_valid. Takes a string representing 7239 forwarded header single value as input and returns bool:TRUE if header is RFC compliant and bool:FALSE otherwise. Example: acl valid req.hdr(forwarded),rfc7239_is_valid #input: "for=127.0.0.1;proto=http" # output: TRUE #input: "proto=custom" # output: FALSE Depends on: - "MINOR: http_ext: introduce http ext converters"	2023-01-27 15:18:59 +01:00
Aurelien DARRAGON	82faad1069	MINOR: http_ext: introduce http ext converters This commit is really simple, it adds the required skeleton code to allow new http_ext converter to be easily registered through STG_REGISTER facility.	2023-01-27 15:18:59 +01:00
Aurelien DARRAGON	f958341610	MINOR: proxy: move 'originalto' option to http_ext Just like forwarded (7239) header and forwardfor header, move parsing, logic and management of 'originalto' option into http_ext dedicated class. We're only doing this to standardize proxy http options management. Existing behavior remains untouched.	2023-01-27 15:18:59 +01:00
Aurelien DARRAGON	730b9836a6	MINOR: proxy: move 'forwardfor' option to http_ext Just like forwarded (7239) header, move parsing, logic and management of 'forwardfor' option into http_ext dedicated class. We're only doing this to standardize proxy http options management. Existing behavior remains untouched.	2023-01-27 15:18:59 +01:00
Aurelien DARRAGON	b2bb9257d2	MINOR: proxy/http_ext: introduce proxy forwarded option Introducing http_ext class for http extension related work that doesn't fit into existing http classes. HTTP extension "forwarded", introduced with 7239 RFC is now supported by haproxy. The option supports various modes from simple to complex usages involving custom sample expressions. Examples : # Those servers want the ip address and protocol of the client request # Resulting header would look like this: # forwarded: proto=http;for=127.0.0.1 backend www_default mode http option forwarded #equivalent to: option forwarded proto for # Those servers want the requested host and hashed client ip address # as well as client source port (you should use seed for xxh32 if ensuring # ip privacy is a concern) # Resulting header would look like this: # forwarded: host="haproxy.org";for="_000000007F2F367E:60138" backend www_host mode http option forwarded host for-expr src,xxh32,hex for_port # Those servers want custom data in host, for and by parameters # Resulting header would look like this: # forwarded: host="host.com";by=_haproxy;for="[::1]:10" backend www_custom mode http option forwarded host-expr str(host.com) by-expr str(_haproxy) for for_port-expr int(10) # Those servers want random 'for' obfuscated identifiers for request # tracing purposes while protecting sensitive IP information # Resulting header would look like this: # forwarded: for=_000000002B1F4D63 backend www_for_hide mode http option forwarded for-expr rand,hex By default (no argument provided), forwarded option will try to mimic x-forward-for common setups (source client ip address + source protocol) The option is not available for frontends. no option forwarded is supported. More info about 7239 RFC here: https://www.rfc-editor.org/rfc/rfc7239.html More info about the feature in doc/configuration.txt This should address feature request GH #575 Depends on: - "MINOR: http_htx: add http_append_header() to append value to header" - "MINOR: sample: add ARGC_OPT" - "MINOR: proxy: introduce http only options"	2023-01-27 15:18:59 +01:00
Aurelien DARRAGON	5f7f5fe76a	MINOR: sample: add ARGC_OPT Add ARGC_OPT enum to provide more context for upcoming sample parse errors involving proxy "option" config directives.	2023-01-27 15:18:59 +01:00
Aurelien DARRAGON	38ebffaf10	MINOR: http_htx: add http_prepend_header() to prepend value to header Just like http_append_header(), but this time to insert new value before an existing one. If the header already contains one or multiple values, ',' is automatically inserted after the new value.	2023-01-27 15:18:59 +01:00
Aurelien DARRAGON	a5a8552cab	MINOR: http_htx: add http_append_header() to append value to header Calling this function as an alternative to http_replace_header_value() to append a new value to existing header instead of replacing the whole header content. If the header already contains one or multiple values: a ',' is automatically appended before the new value. This function is not meant for prepending (providing empty ctx value), in which case we should consider implementing dedicated prepend alternative function.	2023-01-27 15:18:59 +01:00
Willy Tarreau	7cfbb81c85	CLEANUP: mux-h2/trace: shorten the name of the header enc/dec functions The functions in charge of processing headers have their names in the traces and they're among the longest of the mux_h2.c file, while even containing some redundancy. These names are not used outside, let's shorten them: - h2c_decode_headers -> h2c_dec_hdrs - h2s_bck_make_req_headers -> h2s_snd_bhdrs - h2s_frt_make_resp_headers -> h2s_snd_fhdrs Now the traces are a bit more readable: [00\|h2\|5\|mux_h2.c:4822] h2c_dec_hdrs(): h2c=0x1870510(F,FRP) dsi=1 rcvh :method: GET [00\|h2\|5\|mux_h2.c:4822] h2c_dec_hdrs(): h2c=0x1870510(F,FRP) dsi=1 rcvh :path: / [00\|h2\|5\|mux_h2.c:4822] h2c_dec_hdrs(): h2c=0x1870510(F,FRP) dsi=1 rcvh :scheme: http [00\|h2\|5\|mux_h2.c:4822] h2c_dec_hdrs(): h2c=0x1870510(F,FRP) dsi=1 rcvh :authority: localhost:14446 [00\|h2\|5\|mux_h2.c:4822] h2c_dec_hdrs(): h2c=0x1870510(F,FRP) dsi=1 rcvh user-agent: curl/7.54.1 [00\|h2\|5\|mux_h2.c:4822] h2c_dec_hdrs(): h2c=0x1870510(F,FRP) dsi=1 rcvh accept: /	2023-01-26 16:05:51 +01:00
Willy Tarreau	11e8a8c2ac	MEDIUM: mux-h2/trace: add tracing support for headers Now we can make use of TRACE_PRINTF() to iterate over headers as they are received or dumped. It's worth noting that the dumps may occasionally be interrupted due to a buffer full or a realign, but in this case it will be visible because the trace will restart from the first one. All these headers (and trailers) may be interleaved with other connections' so they're all preceeded by the pointer to the connection and optionally the stream (or alternately the stream ID) to help discriminating them. Since it's not easy to read the header directions, sent headers are prefixed with "sndh" and received headers are prefixed with "rcvh", both of which are rare enough in the traces to conveniently support a quick grep. In order to avoid code duplication, h2_encode_headers() was implemented as a wrapper on top of hpack_encode_header(), which optionally emits the header to the trace if the trace is active. In addition, for headers that are encoded using a different method, h2_trace_header() was added as well. Header names are truncated to 256 bytes and values to 1024 bytes. If the lengths are larger, they will be truncated and suffixed with "(... +xxx)" where "xxx" is the number of extra bytes. Example of what an end-to-end H2 request gives: [00\|h2\|5\|mux_h2.c:4818] h2c_decode_headers(): h2c=0x1c13120(F,FRP) dsi=1 rcvh :method: GET [00\|h2\|5\|mux_h2.c:4818] h2c_decode_headers(): h2c=0x1c13120(F,FRP) dsi=1 rcvh :path: / [00\|h2\|5\|mux_h2.c:4818] h2c_decode_headers(): h2c=0x1c13120(F,FRP) dsi=1 rcvh :scheme: http [00\|h2\|5\|mux_h2.c:4818] h2c_decode_headers(): h2c=0x1c13120(F,FRP) dsi=1 rcvh :authority: localhost:14446 [00\|h2\|5\|mux_h2.c:4818] h2c_decode_headers(): h2c=0x1c13120(F,FRP) dsi=1 rcvh user-agent: curl/7.54.1 [00\|h2\|5\|mux_h2.c:4818] h2c_decode_headers(): h2c=0x1c13120(F,FRP) dsi=1 rcvh accept: / [00\|h2\|5\|mux_h2.c:4818] h2c_decode_headers(): h2c=0x1c13120(F,FRP) dsi=1 rcvh cookie: blah [00\|h2\|5\|mux_h2.c:5491] h2s_bck_make_req_headers(): h2c=0x1c1cd90(B,FRH) h2s=0x1c1e3d0(1,IDL) sndh :method: GET [00\|h2\|5\|mux_h2.c:5572] h2s_bck_make_req_headers(): h2c=0x1c1cd90(B,FRH) h2s=0x1c1e3d0(1,IDL) sndh :authority: localhost:14446 [00\|h2\|5\|mux_h2.c:5596] h2s_bck_make_req_headers(): h2c=0x1c1cd90(B,FRH) h2s=0x1c1e3d0(1,IDL) sndh :path: / [00\|h2\|5\|mux_h2.c:5647] h2s_bck_make_req_headers(): h2c=0x1c1cd90(B,FRH) h2s=0x1c1e3d0(1,IDL) sndh user-agent: curl/7.54.1 [00\|h2\|5\|mux_h2.c:5647] h2s_bck_make_req_headers(): h2c=0x1c1cd90(B,FRH) h2s=0x1c1e3d0(1,IDL) sndh accept: / [00\|h2\|5\|mux_h2.c:5647] h2s_bck_make_req_headers(): h2c=0x1c1cd90(B,FRH) h2s=0x1c1e3d0(1,IDL) sndh cookie: blah [00\|h2\|5\|mux_h2.c:4818] h2c_decode_headers(): h2c=0x1c1cd90(B,FRP) dsi=1 rcvh :status: 200 [00\|h2\|5\|mux_h2.c:4818] h2c_decode_headers(): h2c=0x1c1cd90(B,FRP) dsi=1 rcvh content-length: 0 [00\|h2\|5\|mux_h2.c:4818] h2c_decode_headers(): h2c=0x1c1cd90(B,FRP) dsi=1 rcvh x-req: size=102, time=0 ms [00\|h2\|5\|mux_h2.c:4818] h2c_decode_headers(): h2c=0x1c1cd90(B,FRP) dsi=1 rcvh x-rsp: id=dummy, code=200, cache=1, size=0, time=0 ms (0 real) [00\|h2\|5\|mux_h2.c:5210] h2s_frt_make_resp_headers(): h2c=0x1c13120(F,FRH) h2s=0x1c1c780(1,HCR) sndh :status: 200 [00\|h2\|5\|mux_h2.c:5231] h2s_frt_make_resp_headers(): h2c=0x1c13120(F,FRH) h2s=0x1c1c780(1,HCR) sndh content-length: 0 [00\|h2\|5\|mux_h2.c:5231] h2s_frt_make_resp_headers(): h2c=0x1c13120(F,FRH) h2s=0x1c1c780(1,HCR) sndh x-req: size=102, time=0 ms [00\|h2\|5\|mux_h2.c:5231] h2s_frt_make_resp_headers(): h2c=0x1c13120(F,FRH) h2s=0x1c1c780(1,HCR) sndh x-rsp: id=dummy, code=200, cache=1, size=0, time=0 ms (0 real) At some point the frontend/backend names would be useful but that's a more general comment than just the H2 traces.	2023-01-26 15:51:30 +01:00
Willy Tarreau	4b36d5e8de	MINOR: trace: add a trace_no_cb() dummy callback for when to use no callback By default, passing a NULL cb to the trace functions will result in the source's default one to be used. For some cases we won't want to use any callback at all, not event the default one. Let's define a trace_no_cb() function for this, that does absolutely nothing.	2023-01-26 15:49:43 +01:00
Willy Tarreau	8f9a9704bb	MINOR: trace: add a TRACE_ENABLED() macro to determine if a trace is active Sometimes it would be necessary to prepare some messages, pre-process some blocks or maybe duplicate some contents before they vanish for the purpose of tracing them. However we don't want to do that for everything that is submitted to the traces, it's important to do it only for what will really be traced. The __trace() function has all the knowledge for this, to the point of even checking the lockon pointers. This commit splits the function in two, one with the trace decision logic, and the other one for the trace production. The first one is now usable through wrappers such as _trace_enabled() and TRACE_ENABLED() which will indicate whether traces are going to be produced for the current source, level, event mask, parameters and tracking.	2023-01-26 15:49:43 +01:00
Willy Tarreau	80f36b2ac2	CLEANUP: trace: remove the QUIC-specific ifdefs There are ifdefs at several places to only define TRC_ARGS_QCON when QUIC is defined, but nothing prevents this code from building without. Let's just remove those ifdefs, the single "if" they avoid is not worth the extra maintenance burden.	2023-01-26 15:49:43 +01:00
Willy Tarreau	09727ee201	BUG/MINOR: sink: free the forwarding task on exit ASAN reported a small leak of the sink's forwarding task on exit. This should be backported as far as 2.2.	2023-01-26 15:49:32 +01:00
Willy Tarreau	b91910955a	BUG/MINOR: ring: release the backing store name on exit ASAN found that a ring equipped with a backing store did not release the store name on exit. This should be backported to 2.7.	2023-01-26 15:49:31 +01:00
Willy Tarreau	2c701dbc07	BUG/MINOR: log: release global log servers on exit Since 2.6 we have a free_logsrv() function that is used to release log servers. It must be called from deinit() instead of manually iterating over the log servers, otherwise some parts of the structure are not freed (namely the ring name), as reported by ASAN. This should be backported to 2.6.	2023-01-26 15:49:30 +01:00
Willy Tarreau	094ecf19f9	BUG/MEDIUM: hpack: fix incorrect huffman decoding of some control chars Commit `9f4f6b038` ("OPTIM: hpack-huff: reduce the cache footprint of the huffman decoder") replaced the large tables with more space efficient byte arrays, but one table, rht_bit15_11_11_4, has a 64 bytes hole in it that wasn't materialized by filling it with zeroes to make the offsets match, nor by adjusting the offset from the caller. This resulted in some control chars not properly being decoded and being seen as byte 0, and the associated messages to be rejected, as can be seen in issue #1971. This commit fixes it by adjusting the offset used for the higher part of the table so that we don't need to store 64 zeroes that will never be accessed. This needs to be backported to 2.7. Thanks to Christopher for spotting the bug, and to Juanga Covas for providing precious traces showing the problem.	2023-01-26 11:36:39 +01:00
Amaury Denoyelle	b4d119f0c7	BUG/MEDIUM: mux-quic: fix crash on H3 SETTINGS emission A major regression was introduced by following patch commit `71fd03632f` MINOR: mux-quic/h3: send SETTINGS as soon as transport is ready H3 finalize operation is now called at an early stage in the middle of qc_init(). However, some qcc members are not yet initialized. In particular the stream tree which will cause a crash when H3 control stream will be accessed. To fix this, qcc_install_app_ops() has been delayed at the end of qc_init(). This ensures that qcc is properly initialized when app_ops operation are used. This must be backported wherever above patch is. For the record, it has been tagged up to 2.7.	2023-01-25 18:01:18 +01:00
Amaury Denoyelle	19adeb5640	BUG/MINOR: h3: fix GOAWAY emission Since the rework of QUIC streams send scheduling, each stream has to be inserted in QUIC-mux send-list to be able to emit content. This was not the case for GOAWAY which prevent it to be sent. This regression has been introduced by the following patch : commit `20f2a425ff` MAJOR: mux-quic: rework stream sending priorization This new patch fixes the issue by inserting H3 control stream in mux send-list. The impact is deemed minor as for the moment GOAWAY is only sent just before connection/mux cleanup with a CONNECTION_CLOSE. However, it might cause some connections to hang up indefinitely. This should be backported up to 2.7.	2023-01-25 16:09:26 +01:00
Amaury Denoyelle	71fd03632f	MINOR: mux-quic/h3: send SETTINGS as soon as transport is ready As specified by HTTP3 RFC, SETTINGS frame should be sent as soon as possible. Before this patch, this was only done on the first qc_send() invocation. This delay significantly SETTINGS emission until the first H3 response is ready to be transferred. This patch fixes this by ensuring SETTINGS is emitted when MUX-QUIC is being setup. As a side point, return value of finalize operation is checked. This means that an error during SETTINGS emission will cause the connection init to fail. This should be backported up to 2.7.	2023-01-25 16:01:55 +01:00
Olivier Houchard	9a0f8ba837	MINOR: connection: add a BUG_ON() to detect destroying connection in idle list Add a BUG_ON() in conn_free(), to check that when we're freeing a connection, it is not still in the idle connections tree, otherwise the next thread that will try to use it will probably crash.	2023-01-25 15:30:49 +01:00
Remi Tricot-Le Breton	083b230699	MINOR: ssl: Remove debug fprintf in 'update ssl ocsp-response' cli command A debug fprintf was left behind in the new cli function.	2023-01-25 11:51:39 +01:00
Remi Tricot-Le Breton	305a4f32a5	BUG/MINOR: ssl: Fix leaks in 'update ssl ocsp-response' CLI command This patch fixes two leaks in the 'update ssl ocsp-response' cli command. One rather significant one since a whole trash buffer was allocated for every call of the command, and another more marginal one in an error path. This patch does not need to be backported.	2023-01-25 11:51:39 +01:00
Willy Tarreau	fb9a4765b7	BUG/MINOR: sink: make sure to always properly unmap a file-backed ring The munmap() call performed on exit was incorrect since it used to apply to the buffer instead of the area, so neither the pointer nor the size were page-aligned. This patches corrects this and also adds a call to msync() since munmap() alone doesn't guarantee that data will be dumped. This should be backported to 2.6.	2023-01-24 12:11:41 +01:00
Amaury Denoyelle	2d380926ba	MEDIUM: quic-sock: fix udp source address for send on listener socket When receiving a QUIC datagram, destination address is retrieved via recvmsg() and stored in quic-conn as qc.local_addr. This address is then reused when using the quic-conn owned socket. When listener socket mode is preferred, send operation did not specify the source address of the emitted datagram. If listener socket is bound on a wildcard address, the kernel is free to choose any address assigned to the local machine. This may be different from the address selected by the client on its first datagram which will prevent the client to emit next replies. To address this, this patch fixes the UDP source address via sendmsg(). This process is similar to the reception and relies on ancillary message, so the socket is left untouched after the operation. This is heavily platform specific and may not be supported by some kernels. This change has only an impact if listener socket only is used for QUIC communications. This is the default behavior for 2.7 branch but not anymore on 2.8. Use tune.quic.socket-owner set to listener to ensure set it. This should be backported up to 2.7.	2023-01-20 17:06:04 +01:00
Frédéric Lécaille	d18025eeef	BUG/MINOR: quic: Do not request h3 clients to close its unidirection streams It is forbidden to request h3 clients to close its Control and QPACK unidirection streams. If not, the client closes the connection with H3_CLOSED_CRITICAL_STREAM(0x104). Perhaps this could prevent some clients as Chrome to come back for a while. But at quic_conn level there is no mean to identify the streams for which we cannot send STOP_SENDING frame. Such a possibility is even not mentionned in RFC 9000. At this time there is no choice than stopping sending STOP_SENDING frames for all the h3 unidirectional streams inspecting the ->app_opps quic_conn value. Must be backported to 2.7 and 2.6.	2023-01-20 15:49:52 +01:00
Remi Tricot-Le Breton	a0658c3cf3	BUG/MINOR: jwt: Wrong return value checked The wrong return value was checked, resulting in dead code and potential bugs. It should fix GitHub issue #2005. This patch should be backported up to 2.5.	2023-01-20 10:27:37 +01:00
Willy Tarreau	7d84439b48	BUILD: hpack: include global.h for the trash that is needed in debug mode When building with -DDEBUG_HPACK, the trash is needed, but it's declared in global.h. This may be backported to all supported versions.	2023-01-20 00:02:37 +01:00
Willy Tarreau	17c630b846	BUG/MINOR: mux-h2: add missing traces on failed headers decoding In case HPACK cannot be decoded, logs are emitted but there's no info in the H2 traces, so let's add them. This may be backported to all supported versions.	2023-01-20 00:02:21 +01:00
Willy Tarreau	f43f36da5b	BUG/MINOR: mux-h2: make sure to produce a log on invalid requests As reported by Dominik Froehlich in github issue #1968, some H2 request parsing errors do not result in a log being emitted. This is annoying for debugging because while an RST_STREAM is correctly emitted to the client, there's no way without enabling traces to find it on the haproxy side. After some testing with various abnormal requests, a few places were found where logs were missing and could be added. In this case, we simply use sess_log() so some sample fetch functions might not be available since the stream is not created. But at least there will be a BADREQ in the logs. A good eaxmple of this consists in sending forbidden headers or header syntax (e.g. presence of LF in value). Some quick tests can be done this way: - protocol error (LF in value): curl -iv --http2-prior-knowledge -H "$(printf 'a:b\na')" http://0:8001/ - too large header block after decoding: curl -v --http2-prior-knowledge -H "a:$(perl -e "print('a'x10000)")" -H "a:$(perl -e "print('a'x10000)")" http://localhost:8001/ This should be backported where needed, most likely 2.7 and 2.6 at least for a start, and progressively to other versions.	2023-01-19 23:37:00 +01:00
Willy Tarreau	9debe0fb27	BUG/MEDIUM: debug/thread: make the debug handler not wait for !rdv_requests The debug handler may deadlock with some threads waiting for isolation. This may happend during a "show threads" command or even during a panic. The reason is the call to thread_harmless_end() which waits for rdv_requests to turn to zero before releasing its position in thread_dump_state, while that one may not progress if another thread was interrupted in thread_isolate() and is waiting for that thread to drop thread_dump_state. In order to address this, we now use thread_harmless_end_sig() introduced by previous commit: MINOR: threads: add a thread_harmless_end() version that doesn't wait However there's a catch: since commit `f7afdd910` ("MINOR: debug: mark oneself harmless while waiting for threads to finish"), there's a second pair of thread_harmless_now()/thread_harmless_end() that surround the loop around thread_dump_state. Marking a thread harmless before this loop and dropping that without checking rdv_requests there could break the harmless promise made to the other thread if it returns first and proceeds with its isolated work. Hence we just drop this pair which was only preventive for other signal handlers, while as indicated in that patch's commit message, other signals are handled asynchronously and do not require that extra protection. This fix must be backported to 2.7. The problem can be seen by running "show threads" in fast loops (100/s) while reloading haproxy very quickly (10/s) and sending lots of traffic to it (100krps, 15 Gbps). In this case the soft stop calls pool_gc() which isolates a lot and manages to race with the dumps after a few tens of seconds, leaving the process with all threads at 100%.	2023-01-19 19:22:17 +01:00
Willy Tarreau	b2f38c13d1	BUG/MINOR: thread: always reload threads_enabled in loops A few loops waiting for threads to synchronize such as thread_isolate() rightfully filter the thread masks via the threads_enabled field that contains the list of enabled threads. However, it doesn't use an atomic load on it. Before 2.7, the equivalent variables were marked as volatile and were always reloaded. In 2.7 they're fields in ha_tgroup_ctx[], and the risk that the compiler keeps them in a register inside a loop is not null at all. In practice when ha_thread_relax() calls sched_yield() or an x86 PAUSE instruction, it could be verified that the variable is always reloaded. If these are avoided (e.g. architecture providing neither solution), it's visible in asm code that the variables are not reloaded. In this case, if a thread exists just between the moment the two values are read, the loop could spin forever. This patch adds the required _HA_ATOMIC_LOAD() on the relevant threads_enabled fields. It must be backported to 2.7.	2023-01-19 19:22:17 +01:00
Willy Tarreau	ad90110338	BUG/MEDIUM: fd/threads: fix again incorrect thread selection in wakeup broadcast Commit `c1640f79f` ("BUG/MEDIUM: fd/threads: fix incorrect thread selection in wakeup broadcast") fixed an incorrect range being used to pick a thread when broadcasting a wakeup for a foreign thread, but the selection was still wrong as the number of threads and their mask was taken from the current thread instead of the target thread. In addition, the code dealing with the wakeup of a thread from the same group was still relying on MAX_THREADS instead of tg->count. This could theoretically cause random crashes with more than one thread group though this was never encountered. This needs to be backported to 2.7.	2023-01-19 19:22:17 +01:00
Amaury Denoyelle	edfcb55417	MINOR: h3: implement TRAILERS decoding Implement the conversion of H3 request trailers as HTX blocks. This is done through a new function h3_trailers_to_htx(). If the request contains forbidden trailers it is rejected with a stream error. This should be backported up to 2.7.	2023-01-19 16:31:12 +01:00
Christopher Faulet	6bf86c73ba	BUG/MINOR: bwlim: Fix parameters check for set-bandwidth-limit actions First, the inspect-delay is now tested if the action is used on a tcp-response content rule. Then, when an expressions scope is checked, we now take care to detect the right scope depending on the ruleset used (tcp-request, tcp-response, http-request or http-response). This patch could be backported to 2.7.	2023-01-19 16:15:12 +01:00
Christopher Faulet	da2e117369	MEDIUM: bwlim: Support constants limit or period on set-bandwidth-limit actions It is now possible to set a constant for the limit or period parameters on a set-bandwidth-limit actions. The limit must follow the HAProxy size format and is expressed in bytes. The period must follow the HAProxy time format and is expressed in milliseconds. Of course, it is still possible to use sample expressions instead. The documentation was updated accordingly. It is not really a bug. Only exemples were written this way in the documentation. But it could be good to backport this change in 2.7.	2023-01-19 16:15:12 +01:00
Christopher Faulet	ab34ebe5f5	BUG/MINOR: bwlim: Check scope for period expr for set-bandwitdh-limit actions If a period expression is defined for a set-bandwitdh-limit action, its scope must be tested. This patch must be backported to 2.7.	2023-01-19 16:15:12 +01:00
Amaury Denoyelle	4e52010e57	MINOR: h3: implement TRAILERS encoding This patch implement the conversion of an HTX response containing trailer into a H3 HEADERS frame. This is done through a new function named h3_resp_trailers_send(). This was tested with a nginx configuration using <add_trailer> statement. It may be possible that HTX buffer only contains a EOT block without preceeding trailer. In this case, the conversion will produce nothing but fin will be reported. This causes QUIC mux to generate an empty STREAM frame with FIN bit set. This should be backported up to 2.7.	2023-01-19 15:09:01 +01:00
Amaury Denoyelle	7d78eff889	MINOR: h3: extend function for QUIC varint encoding Slighty adjust b_quic_enc_int(). This function is used to encode an integer as a QUIC varint in a struct buffer. A new parameter is added to the function API to specify the width of the encoded integer. By default, 0 should be use to ensure that the minimum space is used. Other valid values are 1, 2, 4 or 8. An error is reported if the width is not large enough. This new parameter will be useful when buffer space is reserved prior to encode an unknown integer value. The maximum size of 8 bytes will be reserved and some data can be put after. When finally encoding the integer, the width can be requested to be 8 bytes. With this new parameter, a small refactoring of the function has been conducted to remove some useless internal variables. This should be backported up to 2.7. It will be mostly useful to implement H3 trailers encoding.	2023-01-19 15:09:01 +01:00
Amaury Denoyelle	8ad2669175	BUG/MINOR: h3: properly handle connection headers Connection headers are not used in HTTP/3. As specified by RFC 9114, a received message containing one of those is considered as malformed and rejected. When converting an HTX message to HTTP/3, these headers are silently skipped. This must be backported up to 2.6. Note that assignment to <h3s.err> must be removed on 2.6 as stream level error has been introduced in 2.7 so this field does not exist in 2.6 A connection error will be used instead automatically.	2023-01-19 15:09:01 +01:00
Willy Tarreau	d1ebee1774	BUG/MINOR: listener: close tiny race between resume_listener() and stopping Pierre Cheynier reported a very rare race condition on soft-stop in the listeners. What happens is that if a previously limited listener is being resumed by another thread finishing an accept loop, and at the same time a soft-stop is performed, the soft-stop will turn the listener's state to LI_INIT, and once the listener's lock is released, resume_listener() in the second thread will try to resume this listener which has an fd==-1, yielding a crash in listener_set_state(): FATAL: bug condition "l->rx.fd == -1" matched at src/listener.c:288 The reason is that resume_listener() only checks for LI_READY, but doesn't consider being called with a non-initialized or a stopped listener. Let's also make sure we don't try to ressuscitate such a listener there. This will have to be backported to all versions.	2023-01-19 11:34:21 +01:00
Remi Tricot-Le Breton	5a8f02ae66	BUG/MEDIUM: jwt: Properly process ecdsa signatures (concatenated R and S params) When the JWT token signature is using ECDSA algorithm (ES256 for instance), the signature is a direct concatenation of the R and S parameters instead of OpenSSL's DER format (see section 3.4 of RFC7518). The code that verified the signatures wrongly assumed that they came in OpenSSL's format and it did not actually work. We now have the extra step of converting the signature into a complete ECDSA_SIG that can be fed into OpenSSL's digest verification functions. The ECDSA signatures in the regtest had to be recalculated and it was made via the PyJWT python library so that we don't end up checking signatures that we built ourselves anymore. This patch should fix GitHub issue #2001. It should be backported up to branch 2.5.	2023-01-18 16:18:31 +01:00
Paul Barnetta	26a9ac5f2f	BUG/MINOR: mux-fcgi: Correctly set pathinfo Existing logic for checking whether a regex subexpression for pathinfo is matched results in valid matches being ignored and non-matches having a new zero length string stored in params->pathinfo. This patch reverses the logic so params->pathinfo is set when the subexpression is matched. Without this patch the example configuration in the documentation: path-info ^(/.+\.php)(/.*)?$ does not result in PATH_INFO being sent to the FastCGI application, as expected, when the second subexpression is matched (in which case both pmatch[2].rm_so and pmatch[2].rm_eo will be non-negative integers). This patch may be backported as far as 2.2, the first release that made the capture of this second subexpression optional.	2023-01-18 07:53:05 +01:00
Fr�d�ric L�caille	21c4c9b854	MINOR: quic: Replace v2 draft definitions by those of the final 2 version This should finalize the support for the QUIC version 2. Must be backported to 2.7.	2023-01-17 16:35:20 +01:00
Fr�d�ric L�caille	33d11c464f	MINOR: sample: Add "quic_enabled" sample fetch This sample fetch returns a boolean. True if the support for QUIC transport protocol was built and if this protocol was not disabled by "no-quic" global option. Must be backported to 2.7.	2023-01-17 16:35:20 +01:00
Fr�d�ric L�caille	12a0317fed	MINOR: quic: Add "no-quic" global option Add "no-quic" to "global" section to disable the use of QUIC transport protocol by all configured QUIC listeners. This is listeners with QUIC addresses on their "bind" lines. Internally, the socket addresses binding is skipped by protocol_bind_all() for receivers with <proto_quic4> or <proto_quic6> as protocol (see protocol struct). Add information about "no-quic" global option to the documentation. Must be backported to 2.7.	2023-01-17 16:35:20 +01:00
Fr�d�ric L�caille	6fc86974cf	MINOR: quic: Disable the active connection migrations Set "disable_active_migration" transport parameter to inform the peer haproxy listeners does not the connection migration feature. Also drop all received datagrams with a modified source address. Must be backported to 2.7.	2023-01-17 16:35:20 +01:00
Fr�d�ric L�caille	f676954f72	MINOR: quic: Useless test about datagram destination addresses There is no reason to check if the peer has modified the destination address of its connection. May be backported to 2.7.	2023-01-17 16:35:20 +01:00
Willy Tarreau	35c4dd0005	CLEANUP: stconn: always use se_fl_set_error() to set the pending error In mux-h2 and mux-quic we still had two places manually setting SE_FL_ERR_PENDING or SE_FL_ERROR depending on the EOS state, instead of using se_fl_set_error() which takes care of the condition. Better use the specialized function for this, it will allow to centralize the conditions. Note that this will be needed to fix a bug.	2023-01-17 16:25:29 +01:00
Willy Tarreau	40725a4eb0	MINOR: listener: also support "quic+" as an address prefix While we do support quic4@ and quic6@ for listening addresses, it was not possible to specify that we want to use an FD inherited from the parent with QUIC. It's just a matter of making it possible to enable a dgram-type socket and a stream-type transport, so let's add this. Now it becomes possible to write "quic+fd@12", "quic+ipv4@addr" etc.	2023-01-16 14:00:51 +01:00
Willy Tarreau	64763342aa	BUG/MINOR: listeners: fix suspend/resume of inherited FDs FDs inherited from a parent process do not deal well with suspend/resume since commit `59b5da487` ("BUG/MEDIUM: listener: never suspend inherited sockets") introduced in 2.3. The problem is that we now report that they cannot be suspended at all, and they return a failure. As such, if a new process fails to bind and sends SIGTTOU to the previous process, that one will notice the failure and instantly switch to soft-stop, leaving no chance to the new process to give up later and signal its failure. What we need to do, however, is to stop receiving new connections from such inherited FDs, which just means that the FD must be unsubscribed from the poller (and resubscribed later if finally it has to stay). With this a new process can start on the already bound FD without problem thanks to the absence of polling, and when the old process stops the new process will be alone on it. This may be backported as far as 2.4.	2023-01-16 14:00:50 +01:00
Willy Tarreau	640e253698	BUG/MINOR: http-ana: make set-status also update txn->status Patrick Hemmer reported an interesting case where the status present in the logs doesn't reflect what was reported to the user. During analysis we could figure that it was in fact solely caused by the code dealing with the set-status action. Indeed, set-status does update the status in the HTX message itself but not in the HTTP transaction. However, at most places where the status is needed to take a decision, it is retrieved from the transaction, and the logs proceed like this as well, though the "status" sample fetch function does retrieve it from the HTX data. This particularly means that once a set-status has been used to modify the status returned to the user, logs do not match that status, and the response code distribution doesn't match either. However a subsequent rule using the status as a condition will still match because the "status" sample fetch function does also extract the status from the HTX stream. Here's an example that fails: frontend f bind :8001 mode http option httplog log stdout daemon http-after-response set-status 400 This will return a 400 to the client but log a 503 and increment http_rsp_5xx. In the end the root cause is that we need to make txn->status the only authoritative place to get the status, and as such it must be updated by the set-status rule. Ideally "status" should just use txn->status but with the two synchronized this way it's not needed. This should be backported since it addresses some consistency issues between logs and what's observed. The set-status action appeared in 1.9 so all stable versions are eligible.	2023-01-13 15:21:08 +01:00
Christopher Faulet	2e47e3a1cf	MINOR: htx: Add an HTX value for the extra field is payload length is unknown When the payload length cannot be determined, the htx extra field is set to the magical vlaue ULLONG_MAX. It is not obvious. This a dedicated HTX value is now used. Now, HTX_UNKOWN_PAYLOAD_LENGTH must be used in this case, instead of ULLONG_MAX.	2023-01-13 11:51:11 +01:00
Christopher Faulet	462f52260c	BUG/MEDIUM: mux-h2: Don't send CANCEL on shutw when response length is unkown Since commit `473e0e54` ("BUG/MINOR: mux-h2: send a CANCEL instead of ES on truncated writes"), a CANCEL may be reported when the response length is unkown. It happens for H1 reponses without "Content-lenght" or "Transfer-encoding" header. Indeed, in this case, the end of the reponse is detected when the server connection is closed. On the fontend side, the H2 multiplexer handles this event as an abort and sensd a RST_STREAM frame with CANCEL error code. The issue is not with the above commit but with the commit `4877045f1` ("MINOR: mux-h2: make streams know if they need to send more data"). The H2_SF_MORE_HTX_DATA flag must only be set if the payload length can be determined. This patch should fix the issue #1992. It must be backported to 2.7.	2023-01-13 11:28:32 +01:00
Christopher Faulet	f2b02cfd94	MAJOR: http-ana: Review error handling during HTTP payload forwarding The error handling in the HTTP payload forwarding is far to be ideal because both sides (request and response) are tested each time. It is espcially ugly on the request side. To report a server error instead of a client error, there are some workarounds to delay the error handling. The reason is that the request analyzer is evaluated before the response one. In addition, errors are tested before the data analysis. It means it is possible to truncate data because errors may be handled to early. So the error handling at this stages was totally reviewed. Aborts are now handled after the data analysis. We also stop to finish the response on request error or the opposite. As a side effect, the HTTP_MSG_ERROR state is now useless. As another side effect, the termination flags are now set by the HTTP analysers and not process_stream().	2023-01-13 11:18:23 +01:00
Christopher Faulet	5aab0a30c5	BUG/MINOR: http-fetch: Don't block HTTP sample fetch eval in HTTP_MSG_ERROR state It was inherited from the legacy HTTP mode, but the message parsing is handled by the underlying mux now. Thus, if a message is in HTTP_MSG_ERROR state, it is just an analysis error and not a parsing error. So there is no reason to block the HTTP sample fetch evaluation in this case. This patch could be backported in all stable versions (For the 2.0, only the htx part must be updated).	2023-01-13 10:58:21 +01:00
Christopher Faulet	f0d80df6e0	MINOR: http-ana: Use http_set_term_flags() when waiting the request body When HAProxy is waiting for the request body and an abort or an error is detected, we can now use http_set_term_flags() function to set the termination flags of the stream instead of handling it by hand.	2023-01-13 10:53:29 +01:00
Christopher Faulet	f4569bbcc1	BUG/MINOR: http-ana: Report SF_FINST_R flag on error waiting the request body When we wait for the request body, we are still in the request analysis. So a SF_FINST_R flag must be reported in logs. Even if some data are already received, at this staged, nothing is sent to the server. This patch could be backported in all stable versions.	2023-01-13 10:49:37 +01:00
Christopher Faulet	4a66c94d25	MINOR: http-ana: Use http_set_term_flags() in most of HTTP analyzers We use the new function to set the HTTP termination flags in the most obvious places. The other places are a bit specific and will be handled one by one in dedicated patched.	2023-01-13 10:24:17 +01:00
Christopher Faulet	71236dedb9	MINOR: http-ana: Add a function to set HTTP termination flags There is already a function to set termination flags but it is not well suited for HTTP streams. So a function, dedicated to the HTTP analysis, was added. This way, this new function will be called for HTTP analysers on error. And if the error is not caugth at this stage, the generic function will still be called from process_stream(). Here, by default a PRXCOND error is reported and depending on the stream state, the reson will be set accordingly: * If the backend SC is in INI state, SF_FINST_T is reported on tarpit and SF_FINST_R otherwise. * SF_FINST_Q is the server connection is queued * SF_FINST_C in any connection attempt state (REQ/TAR/ASS/CONN/CER/RDY). Except for applets, a SF_FINST_R is reported. * Once the server connection is established, SF_FINST_H is reported while HTTP_MSG_DATA state on the response side. * SF_FINST_L is reported if the response is in HTTP_MSG_DONE state or higher and a client error/timeout was reported. * Otherwise SF_FINST_D is reported.	2023-01-13 09:45:23 +01:00
Willy Tarreau	03926129b0	BUG/MEDIUM: peers: make "show peers" more careful about partial initialization Since 2.6 with commit `34e4085f8` ("MEDIUM: peers: Balance applets across threads") the initialization of a peers appctx may be postponed with a wakeup, causing some partially initialized appctx to be visible. The "show peers" command used to only care about peers without appctx, but now it must also take care of those with no stconn, otherwise it can occasionally crash while dumping them. This fix must be backported to 2.6. Thanks to Patrick Hemmer for reporting the problem.	2023-01-12 17:09:34 +01:00
Willy Tarreau	6be8d09a61	OPTIM: global: move byte counts out of global and per-thread During multiple tests we've already noticed that shared stats counters have become a real bottleneck under large thread counts. With QUIC it's pretty visible, with qc_snd_buf() taking 2.5% of the CPU on a 48-thread machine at only 25 Gbps, and this CPU is entirely spent in the atomic increment of the byte count and byte rate. It's also visible in H1/H2 but slightly less since we're working with larger buffers, hence less frequent updates. These counters are exclusively used to report the byte count in "show info" and the byte rate in the stats. Let's move them to the thread_ctx struct and make the stats reader just collect each thread's stats when requested. That's way more efficient than competing on a single cache line. After this, qc_snd_buf has totally disappeared from the perf profile and tests made in h1 show roughly 1% performance increase on small objects.	2023-01-12 16:37:45 +01:00
Remi Tricot-Le Breton	10f113ec55	MINOR: ssl: Reinsert updated ocsp response later in tree in case of http error When updating an OCSP response, in case of HTTP error (host unreachable for instance) we do not want to reinsert the entry at the same place in the update tree otherwise we might retry immediately the update of the same response. This patch adds an arbitrary 1min time to the next_update of a response in such a case. After an HTTP error, instead of waking the update task up after an arbitrary 10s time, we look for the first entry of the update tree and sleep for the apropriate time.	2023-01-12 13:13:45 +01:00
Remi Tricot-Le Breton	1c647adf46	MINOR: ssl: Do not wake ocsp update task if update tree empty In the unlikely event that the ocsp update task is started but the update tree is empty, put the update task to sleep indefinitely. The only way this can happen is if the same certificate is loaded under two different names while the second one has the 'ocsp-update on' option. Since the certificate names are distinct we will have two ckch_stores but a single certificate_ocsp because they are identified by the OCSP_CERTID which is built out of the issuer certificate and the certificate id (which are the same regardless of the .pem file name).	2023-01-12 13:13:45 +01:00
Remi Tricot-Le Breton	474f614975	MINOR: ssl: Treat ocsp-update inconsistencies as fatal errors If incompatibilities are found in a certificate's ocsp-update mode we raised a single alert that will be considered fatal from here on. This is changed because in case of incompatibilities we will end up with an undefined behaviour. The ocsp response might or might not be updated depending on the order in which the multiple ocsp-update options are taken into account.	2023-01-12 13:13:45 +01:00
Remi Tricot-Le Breton	bdd84c5ffb	BUG/MINOR: ssl: OCSP minimum update threshold not properly set An arbitrary 5 minutes minimum interval between two updates of the same OCSP response is defined but it was not properly used when inserting entries in the update tree. This patch does not need to be backported.	2023-01-12 13:13:45 +01:00
Willy Tarreau	145b17fd2f	BUG/MEDIUM: listener: duplicate inherited FDs if needed Since commit `36d9097cf` ("MINOR: fd: Add BUG_ON checks on fd_insert()"), there is currently a test in fd_insert() to detect that we're not trying to reinsert an FD that had already been inserted. This test catches the following anomalies: frontend fail1 bind fd@0 bind fd@0 and: frontend fail2 bind fd@0 shards 2 What happens is that clone_listener() is called on a listener already having an FD, and when sock_{inet,unix}_bind_receiver() are called, the same FD will be registered multiple times and rightfully crash in the sanity check. It wouldn't be correct to block shards though (e.g. they could be used in a default-bind line). What looks like a safer and more future-proof approach simply is to dup() the FD so that each listener has one copy. This is also the only solution that might allow later to support more than 64 threads on an inherited FD. This needs to be backported as far as 2.4. Better wait for at least one extra -dev version before backporting though, as the bug should not be triggered often anyway.	2023-01-11 11:27:20 +01:00
Remi Tricot-Le Breton	8c99081d38	BUG/MINOR: ssl: Missing ssl_conf pointer check when checking ocsp update inconsistencies The ssl_conf might be NULL when processing ocsp_update option in crt-lists. This patch fixes GitHub issue #1995. It does not need to be backported.	2023-01-11 11:20:26 +01:00
Remi Tricot-Le Breton	71237a1457	BUG/MINOR: ssl: Remove unneeded pointer check in ocsp cli release function The ctx pointer cannot be NULL so we can remove the check. This patch fixes GitHub issue #1996. It does not need to be backported.	2023-01-11 11:20:11 +01:00
Christopher Faulet	51dbb4cb79	BUG/MINOR: resolvers: Wait the resolution execution for a do_resolv action The do_resolv action triggers a resolution and must wait for the result. Concretely, if no cache entry is available, it creates a resolution and wakes up the resolvers task. Then it yields. When the action is recalled, if the resolution is still running, it yields again. However, if the resolution is not running, it does not check it was running. Thus, it is possible to ignore the resolution because the action was recalled before the resolvers task had a chance to be executed. If there is result, the action must yield. This patch should fix the issue #1993. It must be backported as far as 2.0.	2023-01-11 10:31:42 +01:00
Christopher Faulet	0ae2e63d85	BUG/MINOR: hlua: Fix Channel.line and Channel.data behavior regarding the doc These both functions are buggy and don't respect the documentation. They must wait for more data, if possible. For Channel.data(), it must happen if not enough data was received orf if no length was specified and no data was received. The first case is properly handled but not the second one. An empty string is return instead. In addition, if there is no data and the channel can't receive more data, 'nil' value must be returned. In the same spirit, for Channel.line(), we must try to wait for more data when no line is found if not enough data was received or if no length was specified. Here again, only the first case is properly handled. And for this function too, 'nil' value must be returned if there is no data and the channel can't receive more data. This patch is related to the issue #1993. It must be backported as far as 2.5.	2023-01-11 10:31:28 +01:00
Christopher Faulet	5f36bfe42e	BUG/MINOR: h1-htx: Remove flags about protocol upgrade on non-101 responses It is possible to have an "upgrade:" header and the corresponding value in the "connection:" header for a non-101 response. It happens for 426-Upgrade-Required messages. However, on HAProxy side, a parsing error is reported for this kind of message because no websocket key header ("sec-websocket-accept:") is found in the response. So a possible fix could be to not perform this test for non-101 responses. However, having flags about protocol upgrade on this kind of response could lead to other bugs. Instead, corresponding flags are removed. Thus, during the H1 response post-parsing, H1_MF_CONN_UPG and H1_MF_UPG_WEBSOCKET flags are removed from any non-101 response. This patch should fix the issue #1997. It must be backported as far as 2.4.	2023-01-11 10:31:28 +01:00
Amaury Denoyelle	a9de7ea1dc	MINOR: mux-quic: use send-list for immediate sending retry Sending is done with several iterations over qcs streams in qc_send(). The first loop is conducted over streams in <qcc.send_list>. After this first iteration, some streams may still have data in their Tx buffer but were blocked by a full qc_stream_desc buffer. In this case, they have release their qc_stream_desc buffer in qcc_streams_sent_done(). New iterations can be done for these streams which can allocate new qc_stream_desc buffer if available. Before this patch, this was done through another stream list <qcc.send_retry_list>. Now, we can reuse the new <qcc.send_list> for this usage. This is safe to use as after first iteration, we have guarantee that either one of the following is true if there is still streams in <qcc.send_list> : * transport layer has rejected data due to congestion * stream is left because it is blocked on stream flow control * stream still has data and has released a fulfilled qc_stream_desc buffer. Immediate retry is useful for these streams : they will allocate a new qc_stream_desc buffer if possible to continue sending. This must be backported up to 2.7.	2023-01-10 18:09:42 +01:00
Amaury Denoyelle	0a1154afb5	MINOR: mux-quic: use send-list for STOP_SENDING/RESET_STREAM emission When a STOP_SENDING or RESET_STREAM must be send, its corresponding qcs is inserted into <qcc.send_list> via qcc_reset_stream() or qcc_abort_stream_read(). This allows to remove the iteration on full qcs tree in qc_send(). Instead, STOP_SENDING and RESET_STREAM is done in the loop over <qcc.send_list> as with STREAM frames. This should improve slightly the performance, most notably when large number of streams are opened. This must be backported up to 2.7.	2023-01-10 17:49:50 +01:00
Amaury Denoyelle	f9b03265f0	MEDIUM: h3: send SETTINGS before STREAM frames Complete qcc_send_stream() function to allow to specify if the stream should be handled in priority. Internally this will insert the qcs instance in front of <qcc.send_list> to be able to treat it before other streams. This functionality is useful when some QUIC streams should be sent before others. Most notably, this is used to guarantee that H3 SETTINGS is done first via the control stream. This must be backported up to 2.7.	2023-01-10 17:49:50 +01:00
Amaury Denoyelle	20f2a425ff	MAJOR: mux-quic: rework stream sending priorization Implement a mechanism to register streams ready to send data in new STREAM frames. Internally, this is implemented with a new list <qcc.send_list> which contains qcs instances. A qcs can be registered safely using the new function qcc_send_stream(). This is done automatically in qc_send_buf() which covers most cases. Also, application layer is free to use it for internal usage streams. This is currently the case for H3 control stream with SETTINGS sending. The main point of this patch is to handle stream sending fairly. This is in stark contrast with previous code where streams with lower ID were always prioritized. This could cause other streams to be indefinitely blocked behind a stream which has a lot of data to transfer. Now, streams are handled in an order scheduled by se_desc layer. This commit is the first one of a serie which will bring other improvments which also relied on the send_list implementation. This must be backported up to 2.7 when deemed sufficiently stable.	2023-01-10 17:49:50 +01:00
Amaury Denoyelle	31d2057c59	MINOR: mux-quic: add traces for flow-control limit reach Add new traces when QUIC flow-control limits are reached at stream or connection level. This may help to explain an interrupted transfer. This should be backported up to 2.6.	2023-01-10 17:45:41 +01:00
Amaury Denoyelle	ab6cdecd71	BUG/MINOR: mux-quic: fix transfer of empty HTTP response QUIC stream did not transferred its response if it was an empty HTTP response without headers nor entity body. This is caused by an incomplete condition on qc_send() which skips streams with empty <tx.buf>. Fix this by extending the condition. Sending will be conducted on a stream if <tx.buf> is not empty or FIN notification must be provided. This allows to send the last STREAM frame for this stream. Such HTTP responses should be extremely rare so this bug is labelled as MINOR. It was encountered with a HTTP/0.9 request on an empty payload. The bug was triggered as HTTP/0.9 does not support header in response message. Also, note that condition to wakeup MUX tasklet has been changed similarly in qc_send_buf(). It is not mandatory to work properly however, most probably because another tasklet_wakeup() is done before/after. This should be backported up to 2.6.	2023-01-10 16:44:53 +01:00
Christopher Faulet	da89e9b95b	MINOR: channel/applets: Stop to test CF_WRITE_ERROR flag if CF_SHUTW is enough In applets, we stop processing when a write error (CF_WRITE_ERROR) or a shutdown for writes (CF_SHUTW) is detected. However, any write error leads to an immediate shutdown for writes. Thus, it is enough to only test if CF_SHUTW is set.	2023-01-09 18:41:08 +01:00
Christopher Faulet	4b490b7517	MINOR: channel: Stop to test CF_READ_ERROR flag if CF_SHUTR is enough When a read error (CF_READ_ERROR) is reported, a shutdown for reads is always performed (CF_SHUTR). Thus, there is no reason to check if CF_READ_ERROR is set if CF_SHUTR is also checked.	2023-01-09 18:41:08 +01:00
Christopher Faulet	2357718217	MEDIUM: channel: Remove CF_READ_ATTACHED and report CF_READ_EVENT instead CF_READ_ATTACHED flag is only used in input events for stream analyzers, CF_MASK_ANALYSER. A read event can be reported instead and this flag can be removed. We must only take care to report a read event when the client connection is upgraded from TCP to HTTP.	2023-01-09 18:41:08 +01:00
Christopher Faulet	049fbcd36a	MINOR: channel: Remove CF_ANA_TIMEOUT and report CF_READ_EVENT instead It appears CF_ANA_TIMEOUT is flag only used in CF_MASK_ANALYSER. All analyzer timeout relies on the analysis expiration date (chn->analyse_exp). Worst, once set, this flag is never removed. Thus this flag can be removed and replaced by a read event (CF_READ_EVENT).	2023-01-09 18:41:08 +01:00
Christopher Faulet	a63f8f379f	MINOR: channel: Remove CF_WRITE_ACTIVITY Thanks to previous changes, CF_WRITE_ACTIVITY flags can be removed. Everywhere it was used, its value is now directly used (CF_WRITE_EVENT\|CF_WRITE_ERROR).	2023-01-09 18:41:08 +01:00
Christopher Faulet	33e03cec5f	MINOR: channel: Remove CF_READ_ACTIVITY Thanks to previous changes, CF_READ_ACTIVITY flags can be removed. Everywhere it was used, its value is now directly used (CF_READ_EVENT\|CF_READ_ERROR).	2023-01-09 18:41:08 +01:00
Christopher Faulet	d898841530	MEDIUM: channel: Use CF_WRITE_EVENT instead of CF_WRITE_PARTIAL Just like CF_READ_PARTIAL, CF_WRITE_PARTIAL is now merged with CF_WRITE_EVENT. There a subtlety in sc_notify(). The "connect" event (formely CF_WRITE_NULL) is now detected with (CF_WRITE_EVENT + sc->state < SC_ST_EST).	2023-01-09 18:41:08 +01:00
Christopher Faulet	285f7616ee	MEDIUM: channel: Use CF_READ_EVENT instead of CF_READ_PARTIAL CF_READ_PARTIAL flag is now merged with CF_READ_EVENT. It means CF_READ_EVENT is set when a read0 is received (formely CF_READ_NULL) or when data are received (formely CF_READ_ACTIVITY). There is nothing special here, except conditions to wake the stream up in sc_notify(). Indeed, the test was a bit changed to reflect recent change. read0 event is now formalized by (CF_READ_EVENT + CF_SHUTR).	2023-01-09 18:41:08 +01:00
Christopher Faulet	b96f2aa380	REORG: channel: Rename CF_WRITE_NULL to CF_WRITE_EVENT As for CF_READ_NULL, it appears CF_WRITE_NULL and other write events on a channel are mainly used to wake up the stream and may be replace by on write event. In this patch, we introduce CF_WRITE_EVENT flag as a replacement to CF_WRITE_EVENT_NULL. There is no breaking change for now, it is just a rename. Gradually, other write events will be merged with this one.	2023-01-09 18:41:08 +01:00
Christopher Faulet	6e1bbc446b	REORG: channel: Rename CF_READ_NULL to CF_READ_EVENT CF_READ_NULL flag is not really useful and used. It is a transient event used to wakeup the stream. As we will see, all read events on a channel may be resumed to only one and are all used to wake up the stream. In this patch, we introduce CF_READ_EVENT flag as a replacement to CF_READ_NULL. There is no breaking change for now, it is just a rename. Gradually, other read events will be merged with this one.	2023-01-09 18:41:08 +01:00
Christopher Faulet	446d8037ce	MINOR: channel: Don't test CF_READ_NULL while CF_SHUTR is enough If CF_READ_NULL flag is set on a channel, it implies a shutdown for reads was performed and CF_SHUTR is also set on this channel. Thus, there is no reason to test is any of these flags is present, testing CF_SHUTR is enough.	2023-01-09 18:41:08 +01:00
Remi Tricot-Le Breton	14419ebf2b	MINOR: ssl: Remove mention of ckch_store in error message of cli command When calling 'update ssl ocsp-response' with an unknown certificate file name, the error message would mention a "ckch_store" which is an internal structure unknown by users.	2023-01-09 15:43:41 +01:00
Remi Tricot-Le Breton	648c83ecdd	MINOR: ssl: Limit ocsp_uri buffer size to minimum The ocsp_uri field of the certificate_ocsp structure was a 16k buffer when it could be hand allocated to just the required size to store the OCSP uri. This field is now behaving the same way as the sctl and ocsp_response buffers of the ckch_store structure.	2023-01-09 15:43:41 +01:00
Remi Tricot-Le Breton	2d1daa8095	BUG/MINOR: ssl: Fix OCSP_CERTID leak when same certificate is used multiple times If a given certificate is used multiple times in a configuration, the ocsp_cid field would have been overwritten during each ssl_sock_load_ocsp call even if it was previously filled. This patch does not need to be backported.	2023-01-09 15:43:41 +01:00
Remi Tricot-Le Breton	fc92b8bda5	MINOR: ssl: Detect more OCSP update inconsistencies If a configuration such as the following was included in a crt-list file, it would not have raised a warning about 'ocsp-update' inconsistencies for the concerned certificate: cert.pem [ocsp-update on] cert.pem because the second line as a NULL entry->ssl_conf.	2023-01-09 15:43:41 +01:00
Remi Tricot-Le Breton	14d7f0eb48	MINOR: ssl: Release ssl_ocsp_task_ctx.cur_ocsp when destroying task In the unlikely event that the OCSP udpate task is killed in the middle of an update process (request sent but no response received yet) the cur_ocsp member of the update context would keep an unneeded reference to a certificate_ocsp object. It must then be freed during the task's cleanup.	2023-01-09 15:43:41 +01:00
Remi Tricot-Le Breton	112b16a4d0	MINOR: ssl: Only set ocsp->issuer if issuer not in cert chain If the ocsp issuer certificate was actually taken from the certificate chain in ssl_sock_load_ocsp, we don't need to keep an extra reference on it since we already keep a reference to the full certificate chain.	2023-01-09 15:43:41 +01:00
Remi Tricot-Le Breton	8bdd0050e2	MINOR: ssl: Create temp X509_STORE filled with cert chain when checking ocsp response When calling OCSP_basic_verify to check the validity of the received OCSP response, we need to provide an untrusted certificate chain as well as an X509_STORE holding only trusted certificates. Since the certificate chain and the issuer certificate are all provided by the user, we assume that they are valid and we add them all to a temporary store. This enables to focus only on the response's validity.	2023-01-09 15:43:41 +01:00
Remi Tricot-Le Breton	57f60c2316	BUG/MINOR: ssl: Crash during cleanup because of ocsp structure pointer UAF When ocsp-update is enabled for a given certificate, its certificate_ocsp objects is inserted in two separate trees (the actual ocsp response one and the ocsp update one). But since the same instance is used for the two trees, its ownership is kept by the regular ocsp response one. The ocsp update task should then never have to free the ocsp entries. The crash actually occurred because of this. The update task was freeing entries whose reference counter was not increased while a reference was still held by the SSL_CTXs. The only time during which the ocsp update task will need to increase the reference counter is during an actual update, because at this moment the entry is taken out of the update tree and a 'flying' reference to the certificate_ocsp is kept in the ocsp update context. This bug could be reproduced by calling './haproxy -f conf.cfg -c' with any of the used certificates having the 'ocsp-update on' option. For some reason asan caught the bug easily but valgrind did not. This patch does not need to be backported.	2023-01-09 15:43:41 +01:00
Remi Tricot-Le Breton	15dc0e2a1c	BUG/MINOR: ssl: Fix crash in 'update ssl ocsp-response' CLI command This CLI command crashed when called for a certificate which did not have an OCSP response during startup because it assumed that the ocsp_issuer pointer of the ckch_data object would be valid. It was only true for already known OCSP responses though. The ocsp issuer certificate is now taken either from the ocsp_issuer pointer or looked for in the certificate chain. This is the same logic as the one in ssl_sock_load_ocsp. This patch does not need to be backported.	2023-01-09 15:43:41 +01:00
Manu Nicolas	45b6b23335	CLEANUP: htx: fix a typo in an error message of http_str_to_htx This fixes a typo in an error message about headers in the http_str_to_htx function.	2023-01-09 05:28:03 +01:00
Willy Tarreau	40c88f997f	[RELEASE] Released version 2.8-dev1 Released version 2.8-dev1 with the following main changes : - MEDIUM: 51d: add support for 51Degrees V4 with Hash algorithm - MINOR: debug: support pool filtering on "debug dev memstats" - MINOR: debug: add a balance of alloc - free at the end of the memstats dump - LICENSE: wurfl: clarify the dummy library license. - MINOR: event_hdl: add event handler base api - DOC/MINOR: api: add documentation for event_hdl feature - MEDIUM: ssl: rename the struct "cert_key_and_chain" to "ckch_data" - MINOR: quic: remove qc from quic_rx_packet - MINOR: quic: complete traces in qc_rx_pkt_handle() - MINOR: quic: extract datagram parsing code - MINOR: tools: add port for ipcmp as optional criteria - MINOR: quic: detect connection migration - MINOR: quic: ignore address migration during handshake - MINOR: quic: startup detect for quic-conn owned socket support - MINOR: quic: test IP_PKTINFO support for quic-conn owned socket - MINOR: quic: define config option for socket per conn - MINOR: quic: allocate a socket per quic-conn - MINOR: quic: use connection socket for emission - MEDIUM: quic: use quic-conn socket for reception - MEDIUM: quic: move receive out of FD handler to quic-conn io-cb - MINOR: mux-quic: rename duplicate function names - MEDIUM: quic: requeue datagrams received on wrong socket - MINOR: quic: reconnect quic-conn socket on address migration - MINOR: quic: activate socket per conn by default - BUG/MINOR: ssl: initialize SSL error before parsing - BUG/MINOR: ssl: initialize WolfSSL before parsing - BUG/MINOR: quic: fix fd leak on startup check quic-conn owned socket - BUG/MEDIIM: stconn: Flush output data before forwarding close to write side - MINOR: server: add srv->rid (revision id) value - MINOR: stats: add server revision id support - MINOR: server/event_hdl: add support for SERVER_ADD and SERVER_DEL events - MINOR: server/event_hdl: add support for SERVER_UP and SERVER_DOWN events - BUG/MEDIUM: checks: do not reschedule a possibly running task on state change - BUG/MINOR: checks: make sure fastinter is used even on forced transitions - CLEANUP: assorted typo fixes in the code and comments - MINOR: mworker: display an alert upon a wait-mode exit - BUG/MEDIUM: mworker: fix segv in early failure of mworker mode with peers - BUG/MEDIUM: mworker: create the mcli_reload socketpairs in case of upgrade - BUG/MINOR: checks: restore legacy on-error fastinter behavior - MINOR: check: use atomic for s->consecutive_errors - MINOR: stats: properly handle ST_F_CHECK_DURATION metric - MINOR: mworker: remove unused legacy code in mworker_cleanlisteners - MINOR: peers: unused code path in process_peer_sync - BUG/MINOR: init/threads: continue to limit default thread count to max per group - CLEANUP: init: remove useless assignment of nbthread - BUILD: atomic: atomic.h may need compiler.h on ARMv8.2-a - BUILD: makefile/da: also clean Os/ in Device Atlas dummy lib dir - BUG/MEDIUM: httpclient/lua: double LIST_DELETE on end of lua task - CLEANUP: pools: move the write before free to the uaf-only function - CLEANUP: pool: only include pool-os from pool.c not pool.h - REORG: pool: move all the OS specific code to pool-os.h - CLEANUP: pools: get rid of CONFIG_HAP_POOLS - DEBUG: pool: show a few examples in -dMhelp - MINOR: pools: make DEBUG_UAF a runtime setting - BUG/MINOR: promex: create haproxy_backend_agg_server_status - MINOR: promex: introduce haproxy_backend_agg_check_status - DOC: promex: Add missing backend metrics - BUG/MAJOR: fcgi: Fix uninitialized reserved bytes - REGTESTS: fix the race conditions in iff.vtc - CI: github: reintroduce openssl 1.1.1 - BUG/MINOR: quic: properly handle alloc failure in qc_new_conn() - BUG/MINOR: quic: handle alloc failure on qc_new_conn() for owned socket - CLEANUP: mux-quic: remove unused attribute on qcs_is_close_remote() - BUG/MINOR: mux-quic: remove qcs from opening-list on free - BUG/MINOR: mux-quic: handle properly alloc error in qcs_new() - CI: github: split ssl lib selection based on git branch - REGTESTS: startup: check maxconn computation - BUG/MINOR: startup: don't use internal proxies to compute the maxconn - REGTESTS: startup: change the expected maxconn to 11000 - CI: github: set ulimit -n to a greater value - REGTESTS: startup: activate automatic_maxconn.vtc - MINOR: sample: add param converter - CLEANUP: ssl: remove check on srv->proxy - BUG/MEDIUM: freq-ctr: Don't compute overshoot value for empty counters - BUG/MEDIUM: resolvers: Use tick_first() to update the resolvers task timeout - REGTESTS: startup: add alternatives values in automatic_maxconn.vtc - BUG/MEDIUM: h3: reject request with invalid header name - BUG/MEDIUM: h3: reject request with invalid pseudo header - MINOR: http: extract content-length parsing from H2 - BUG/MEDIUM: h3: parse content-length and reject invalid messages - CI: github: remove redundant ASAN loop - CI: github: split matrix for development and stable branches - BUG/MEDIUM: mux-h1: Don't release H1 stream upgraded from TCP on error - BUG/MINOR: mux-h1: Fix test instead a BUG_ON() in h1_send_error() - MINOR: http-htx: add BUG_ON to prevent API error on http_cookie_register - BUG/MEDIUM: h3: fix cookie header parsing - BUG/MINOR: h3: fix memleak on HEADERS parsing failure - MINOR: h3: check return values of htx_add_* on headers parsing - MINOR: ssl: Remove unneeded buffer allocation in show ocsp-response - MINOR: ssl: Remove unnecessary alloc'ed trash chunk in show ocsp-response - BUG/MINOR: ssl: Fix memory leak of find_chain in ssl_sock_load_cert_chain - MINOR: stats: provide ctx for dumping functions - MINOR: stats: introduce stats field ctx - BUG/MINOR: stats: fix show stat json buffer limitation - MINOR: stats: make show info json future-proof - BUG/MINOR: quic: fix crash on PTO rearm if anti-amplification reset - BUILD: 51d: fix build issue with recent compilers - REGTESTS: startup: disable automatic_maxconn.vtc - BUILD: peers: peers-t.h depends on stick-table-t.h - BUG/MEDIUM: tests: use tmpdir to create UNIX socket - BUG/MINOR: mux-h1: Report EOS on parsing/internal error for not running stream - BUG/MINOR:: mux-h1: Never handle error at mux level for running connection - BUG/MEDIUM: stats: Rely on a local trash buffer to dump the stats - OPTIM: pool: split the read_mostly from read_write parts in pool_head - MINOR: pool: make the thread-local hot cache size configurable - MINOR: freq_ctr: add opportunistic versions of swrate_add() - MINOR: pool: only use opportunistic versions of the swrate_add() functions - REGTESTS: ssl: enable the ssl_reuse.vtc test for WolfSSL - BUG/MEDIUM: mux-quic: fix double delete from qcc.opening_list - BUG/MEDIUM: quic: properly take shards into account on bind lines - BUG/MINOR: quic: do not allocate more rxbufs than necessary - MINOR: ssl: Add a lock to the OCSP response tree - MINOR: httpclient: Make the CLI flags public for future use - MINOR: ssl: Add helper function that extracts an OCSP URI from a certificate - MINOR: ssl: Add OCSP request helper function - MINOR: ssl: Add helper function that checks the validity of an OCSP response - MINOR: ssl: Add "update ssl ocsp-response" cli command - MEDIUM: ssl: Add ocsp_certid in ckch structure and discard ocsp buffer early - MINOR: ssl: Add ocsp_update_tree and helper functions - MINOR: ssl: Add crt-list ocsp-update option - MINOR: ssl: Store 'ocsp-update' mode in the ckch_data and check for inconsistencies - MEDIUM: ssl: Insert ocsp responses in update tree when needed - MEDIUM: ssl: Add ocsp update task main function - MEDIUM: ssl: Start update task if at least one ocsp-update option is set to on - DOC: ssl: Add documentation for ocsp-update option - REGTESTS: ssl: Add tests for ocsp auto update mechanism - MINOR: ssl: Move OCSP code to a dedicated source file - BUG/MINOR: ssl/ocsp: check chunk_strcpy() in ssl_ocsp_get_uri_from_cert() - CLEANUP: ssl/ocsp: add spaces around operators - BUG/MEDIUM: mux-h2: Refuse interim responses with end-stream flag set - BUG/MINOR: pool/stats: Use ullong to report total pool usage in bytes in stats - BUG/MINOR: ssl/ocsp: httpclient blocked when doing a GET - MINOR: httpclient: don't add body when istlen is empty - MEDIUM: httpclient: change the default log format to skip duplicate proxy data - BUG/MINOR: httpclient/log: free of invalid ptr with httpclient_log_format - MEDIUM: mux-quic: implement shutw - MINOR: mux-quic: do not count stream flow-control if already closed - MINOR: mux-quic: handle RESET_STREAM reception - MEDIUM: mux-quic: implement STOP_SENDING emission - MINOR: h3: use stream error when needed instead of connection - CI: github: enable github api authentication for OpenSSL tags read - BUG/MINOR: mux-quic: ignore remote unidirectional stream close - CI: github: use the GITHUB_TOKEN instead of a manually generated token - BUILD: makefile: build the features list dynamically - BUILD: makefile: move common options-oriented macros to include/make/options.mk - BUILD: makefile: sort the features list - BUILD: makefile: initialize all build options' variables at once - BUILD: makefile: add a function to collect all options' CFLAGS/LDFLAGS - BUILD: makefile: start to automatically collect CFLAGS/LDFLAGS - BUILD: makefile: ensure that all USE_* handlers appear before CFLAGS are used - BUILD: makefile: clean the wolfssl include and lib generation rules - BUILD: makefile: make sure to also ignore SSL_INC when using wolfssl - BUILD: makefile: reference libdl only once - BUILD: makefile: make sure LUA_INC and LUA_LIB are always initialized - BUILD: makefile: do not restrict Lua's prepend path to empty LUA_LIB_NAME - BUILD: makefile: never force -latomic, set USE_LIBATOMIC instead - BUILD: makefile: add an implicit USE_MATH variable for -lm - BUILD: makefile: properly report USE_PCRE/USE_PCRE2 in features - CLEANUP: makefile: properly indent ifeq/ifneq conditional blocks - BUILD: makefile: rework 51D to split v3/v4 - BUILD: makefile: support LIBCRYPT_LDFLAGS - BUILD: makefile: support RT_LDFLAGS - BUILD: makefile: support THREAD_LDFLAGS - BUILD: makefile: support BACKTRACE_LDFLAGS - BUILD: makefile: support SYSTEMD_LDFLAGS - BUILD: makefile: support ZLIB_CFLAGS and ZLIB_LDFLAGS - BUILD: makefile: support ENGINE_CFLAGS - BUILD: makefile: support OPENSSL_CFLAGS and OPENSSL_LDFLAGS - BUILD: makefile: support WOLFSSL_CFLAGS and WOLFSSL_LDFLAGS - BUILD: makefile: support LUA_CFLAGS and LUA_LDFLAGS - BUILD: makefile: support DEVICEATLAS_CFLAGS and DEVICEATLAS_LDFLAGS - BUILD: makefile: support PCRE[2]_CFLAGS and PCRE[2]_LDFLAGS - BUILD: makefile: refactor support for 51DEGREES v3/v4 - BUILD: makefile: support WURFL_CFLAGS and WURFL_LDFLAGS - BUILD: makefile: make all OpenSSL variants use the same settings - BUILD: makefile: remove the special case of the SSL option - BUILD: makefile: only consider settings from enabled options - BUILD: makefile: also list per-option settings in 'make opts' - BUG/MINOR: debug: don't mask the TH_FL_STUCK flag before dumping threads - MINOR: cfgparse-ssl: avoid a possible crash on OOM in ssl_bind_parse_npn() - BUG/MINOR: ssl: Missing goto in error path in ocsp update code - BUG/MINOR: stick-table: report the correct action name in error message - CI: Improve headline in matrix.py - CI: Add in-memory cache for the latest OpenSSL/LibreSSL - CI: Use proper `if` blocks instead of conditional expressions in matrix.py - CI: Unify the `GITHUB_TOKEN` name across matrix.py and vtest.yml - CI: Explicitly check environment variable against `None` in matrix.py - CI: Reformat `matrix.py` using `black` - MINOR: config: add environment variables for default log format - REGTESTS: Remove REQUIRE_VERSION=1.9 from all tests - REGTESTS: Remove REQUIRE_VERSION=2.0 from all tests - REGTESTS: Remove tests with REQUIRE_VERSION_BELOW=1.9 - BUG/MINOR: http-fetch: Only fill txn status during prefetch if not already set - BUG/MAJOR: buf: Fix copy of wrapping output data when a buffer is realigned - DOC: config: fix alphabetical ordering of http-after-response rules - MINOR: http-rules: Add missing actions in http-after-response ruleset - DOC: config: remove duplicated "http-response sc-set-gpt0" directive - BUG/MINOR: proxy: free orgto_hdr_name in free_proxy() - REGTEST: fix the race conditions in json_query.vtc - REGTEST: fix the race conditions in add_item.vtc - REGTEST: fix the race conditions in digest.vtc - REGTEST: fix the race conditions in hmac.vtc - BUG/MINOR: fd: avoid bad tgid assertion in fd_delete() from deinit() - BUG/MINOR: http: Memory leak of http redirect rules' format string - MEDIUM: stick-table: set the track-sc limit at boottime via tune.stick-counters - MINOR: stick-table: implement the sc-add-gpc() action	2023-01-07 09:45:17 +01:00
Willy Tarreau	5a72d03a58	MINOR: stick-table: implement the sc-add-gpc() action This action increments the General Purpose Counter at the index <idx> of the array associated to the sticky counter designated by <sc-id> by the value of either integer <int> or the integer evaluation of expression <expr>. Integers and expressions are limited to unsigned 32-bit values. If an error occurs, this action silently fails and the actions evaluation continues. <idx> is an integer between 0 and 99 and <sc-id> is an integer between 0 and 2. It also silently fails if the there is no GPC stored at this index. The entry in the table is refreshed even if the value is zero. The 'gpc_rate' is automatically adjusted to reflect the average growth rate of the gpc value. The main use of this action is to count scores or total volumes (e.g. estimated danger per source IP reported by the server or a WAF, total uploaded bytes, etc).	2023-01-07 09:11:22 +01:00
Willy Tarreau	6c0117168e	MEDIUM: stick-table: set the track-sc limit at boottime via tune.stick-counters The number of stick-counter entries usable by track-sc rules is currently set at build time. There is no good value for this since the vast majority of users don't need any, most need only a few and rare users need more. Adding more counters for everyone increases memory and CPU usages for no reason. This patch moves the per-session and per-stream arrays to a pool of a size defined at boot time. This way it becomes possible to set the number of entries at boot time via a new global setting "tune.stick-counters" that sets the limit for the whole process. When not set, the MAX_SESS_STR_CTR value still applies, or 3 if not set, as before. It is also possible to lower the value to 0 to save a bit of memory if not used at all. Note that a few low-level sample-fetch functions had to be protected due to the ability to use sample-fetches in the global section to set some variables.	2023-01-06 18:08:49 +01:00
Remi Tricot-Le Breton	3120284c29	BUG/MINOR: http: Memory leak of http redirect rules' format string When the configuration contains such a line: http-request redirect location / a "struct logformat_node" object is created and it contains an "arg" member which gets alloc'ed as well in which we copy the new location (see add_to_logformat_list). This internal arg pointer was not freed in the dedicated release_http_redir release function. Likewise, the expression pointer was not released as well. This patch can be backported to all stable branches. It should apply as-is all the way to 2.2 but it won't on 2.0 because release_http_redir did not exist yet.	2023-01-06 16:42:24 +01:00
Willy Tarreau	80ff10c81d	BUG/MINOR: fd: avoid bad tgid assertion in fd_delete() from deinit() In 2.7, commit `0dc1cc93b` ("MAJOR: fd: grab the tgid before manipulating running") added a check to make sure we never try to delete an FD from the wrong thread group. It already handles the specific case of an isolated thread (e.g. stop a listener from the CLI) but forgot to take into account the deinit() code iterating over all idle server connections to close them. This results in the crash below during deinit() if thread groups are enabled and idle connections exist on a thread group higher than 1. [WARNING] (15711) : Proxy decrypt stopped (cumulated conns: FE: 64, BE: 374511). [WARNING] (15711) : Proxy stats stopped (cumulated conns: FE: 0, BE: 0). [WARNING] (15711) : Proxy GLOBAL stopped (cumulated conns: FE: 0, BE: 0). FATAL: bug condition "fd_tgid(fd) != ti->tgid && !thread_isolated()" matched at src/fd.c:369 call trace(11): \| 0x4a6060 [c6 04 25 01 00 00 00 00]: main-0x1d60 \| 0x67fcc6 [c7 43 68 fd ad de fd 5b]: sock_conn_ctrl_close+0x16/0x1f \| 0x59e6f5 [48 89 ef e8 83 65 11 00]: main+0xf6935 \| 0x60ad16 [48 8b 1b 48 81 fb a0 91]: free_proxy+0x716/0xb35 \| 0x62750e [48 85 db 74 35 48 89 dd]: deinit+0xbe/0x87a \| 0x627ce2 [89 ef e8 97 76 e7 ff 0f]: deinit_and_exit+0x12/0x19 \| 0x4a9694 [bf e6 ff 9d 00 44 89 6c]: main+0x18d4/0x2c1a There's no harm though since all traffic already ended. This must be backported to 2.7.	2023-01-05 18:06:58 +01:00
Aurelien DARRAGON	b2d797a53f	BUG/MINOR: proxy: free orgto_hdr_name in free_proxy() Unlike fwdfor_hdr_name, orgto_hdr_name was not properly freed in free_proxy(). This did not cause observable memory leaks because originalto proxy option is only used for user configurable proxies, which are solely freed right before process termination. No backport needed unless some architectural changes causing regular proxies to be freed and reused multiple times within single process lifetime are made.	2023-01-05 15:20:41 +01:00
Christopher Faulet	a92480462c	MINOR: http-rules: Add missing actions in http-after-response ruleset This patch adds the support of following actions in the http-after-response ruleset: * set-map, del-map and del-acl * set-log-level * sc-inc-gpc, sc-inc-gpc0 and set-inc-gpc1 * sc-inc-gpt and sc-set-gpt0 This patch should solve the issue #1980.	2023-01-05 11:23:59 +01:00
Christopher Faulet	31850b470a	BUG/MINOR: http-fetch: Only fill txn status during prefetch if not already set When an HTTP sample fetch is evaluated, a prefetch is performed to check the channel contains a valid HTTP message. If the HTTP analysis was not already started, some info are filled. It may be an issue when an error is returned before the response analysis and when http-after-response rules are used because the original HTTP txn status may be crushed. For instance, with the following configuration: listen l1 log global mode http bind :8000 log-format ST=%ST http-after-response set-status 400 #http-after-response set-var(res.foo) status A "ST=503" is reported in the log messages, independantly on the first http-after-response rule. The same must happen if the second rule is uncommented. However, for now, a "ST=400" is logged. To fix the bug, during the prefetch, the HTTP txn status is only set if it is undefined (-1). This way, we are sure the original one is never lost. This patch should be backported as far as 2.2.	2023-01-05 09:33:23 +01:00
S�bastien Gross	537b9e7f36	MINOR: config: add environment variables for default log format This patch provides a convenient way to override the default TCP, HTTP and HTTP log formats. Instead of having a look into the documentation to figure out what is the appropriate default log format three new environment variables can be used: HAPROXY_TCP_LOG_FMT, HAPROXY_HTTP_LOG_FMT and HAPROXY_HTTPS_LOG_FMT. Their content are substituted verbatim. These variables are set before parsing the configuration and are unset just after all configuration files are successful parsed. Example: # Instead of writing this long log-format line... log-format "%ci:%cp [%tr] %ft %b/%s %TR/%Tw/%Tc/%Tr/%Ta %ST %B %CC \ %CS %tsc %ac/%fc/%bc/%sc/%rc %sq/%bq %hr %hs %{+Q}r \ lr=last_rule_file:last_rule_line" # ..the HAPROXY_HTTP_LOG_FMT can be used to provide the default # http log-format string log-format "${HAPROXY_HTTP_LOG_FMT} lr=last_rule_file:last_rule_line" Please note that nothing prevents users to unset the variables or override their content in a global section. Signed-off-by: S�bastien Gross <sgross@haproxy.com>	2023-01-04 08:23:43 +01:00
Willy Tarreau	20391519c3	BUG/MINOR: stick-table: report the correct action name in error message sc-inc-gpc() learned to use arrays in 2.5 with commit `4d7ada8f9` ("MEDIUM: stick-table: add the new arrays of gpc and gpc_rate"), but the error message says "sc-set-gpc" instead of "sc-inc-gpc". Let's fix this to avoid confusion. This can be backported to 2.5.	2023-01-02 17:35:50 +01:00
Remi Tricot-Le Breton	c389b04bc5	BUG/MINOR: ssl: Missing goto in error path in ocsp update code When converting an OCSP request's information into base64, the return value of a2base64 is checked but processing is not interrupted when it returns a negative value, which was caught by coverity. This patch fixes GitHub issue #1974. It does not need to be backported.	2023-01-02 15:21:57 +01:00
Willy Tarreau	c57fb3be75	MINOR: cfgparse-ssl: avoid a possible crash on OOM in ssl_bind_parse_npn() Upon out of memory condition at boot, we could possibly crash when parsing the "npn" bind line keyword since it's used unchecked. There's no real need to backport this though it will not hurt.	2023-01-02 09:51:35 +01:00
Willy Tarreau	b5662519df	BUG/MINOR: debug: don't mask the TH_FL_STUCK flag before dumping threads Commit `f0c86ddfe` ("BUG/MEDIUM: debug: fix parallel thread dumps again") added a clearing of the TH_FL_STUCK flag before dumping threads in case of parallel dumps, but that was in part a sort of workaround for some remains of the commit that introduced the flag in 2.0 before the watchdog existed, and which would set it after dumping a thread: `e6a02fa65` ("MINOR: threads: add a "stuck" flag to the thread_info struct"), and in part an attempt to avoid that a thread waiting for too long during the dump would get the flag set. But that is not possible, a thread waiting for being dumped has the harmless bit set and doesn't get the stuck bit. What happens in fact is that issuing "show threads" in fast loops ends up causing some threads to keep their STUCK bit that was set at the end of "show threads", and confuses the output. The problem with doing this is that the flag is cleared before the thread is dumped, and since this flag is used to decide whether to show a backtrace or not, we don't get backtraces anymore of stuck threads since the commit above in 2.7. This patch just removes the two points where the flag was cleared by the commit above. It should be backported to 2.7.	2023-01-02 09:51:35 +01:00
Amaury Denoyelle	9107731358	BUG/MINOR: mux-quic: ignore remote unidirectional stream close Remove ABORT_NOW() on remote unidirectional stream closure. This is required to ensure our implementation is evolutive enough to not fail on unknown stream type. Note that for the moment MAX_STREAMS_UNI flow-control frame is never emitted. This should be unnecessary for HTTP/3 which have a limited usage of unidirectional streams but may be required if other application protocols are supported in the future. ABORT_NOW() was triggered by s2n-quic which opens an unknown unidirectional stream with greasing. This was detected by QUIC interop runner for http3 testcase. This must be backported up to 2.6.	2022-12-23 00:15:20 +01:00
Amaury Denoyelle	2fe93ab2d7	MINOR: h3: use stream error when needed instead of connection Use a stream error when possible instead of always closing the whole connection. This requires a new field <err> in h3s structure. Change slightly the decoding loop to facilitate error propagation. It will be interrupted as soon as <h3s.err> or <h3c.err> is non null. In the later case, a CONNECTION_CLOSE is requested through qcc_emit_cc_app(). For stream error, H3 layer uses qcc_abort_stream_read() coupled with qcc_reset_stream(). This is in conformance with RFC 9114 which recommends to use STOP_SENDING + RESET_STREAM emission on stream error. This commit is part of implementing H3 errors at the stream level. This should be backported up to 2.7.	2022-12-22 16:47:24 +01:00
Amaury Denoyelle	663e872e3a	MEDIUM: mux-quic: implement STOP_SENDING emission Implement STOP_SENDING. This is divided in two main functions : * qcc_abort_stream_read() which can be used by application protocol to request for a STOP_SENDING. This set the flag QC_SF_READ_ABORTED. * qcs_send_reset() is a static function called after the preceding one. It will send a STOP_SENDING via qcc_send(). QC_SF_READ_ABORTED flag is now properly used : if activated on a stream during qcc_recv(), <qcc.app_ops.decode_qcs> callback is skipped. Also, abort reading on unknown unidirection remote stream is now fully supported with the emission of a STOP_SENDING as specified by RFC 9000. This commit is part of implementing H3 errors at the stream level. This will allows the H3 layer to request the peer to close its endpoint for an error on a stream. This should be backported up to 2.7.	2022-12-22 16:38:16 +01:00
Amaury Denoyelle	5854fc08cc	MINOR: mux-quic: handle RESET_STREAM reception Implement RESET_STREAM reception by mux-quic. On reception, qcs instance will be mark as remotely closed and its Rx buffer released. The stream layer will be flagged on error if still attached. This commit is part of implementing H3 errors at the stream level. Indeed, on H3 stream errors, STOP_SENDING + RESET_STREAM should be emitted. The STOP_SENDING will in turn generate a RESET_STREAM by the remote peer which will be handled thanks to this patch. This should be backported up to 2.7.	2022-12-22 16:38:04 +01:00
Amaury Denoyelle	bb6296ce06	MINOR: mux-quic: do not count stream flow-control if already closed It is unnecessary to increase stream credit once its size is known. Indeed, a peer cannot sent a greater offset than the value advertized. Else, connection will be closed on STREAM reception with FINAL_SIZE_ERROR. This commit is a small optimization and may prevent the emission of unneeded MAX_STREAM_DATA frames on some occasions. It should be backported up to 2.7.	2022-12-22 16:29:59 +01:00
Amaury Denoyelle	a473f196f1	MEDIUM: mux-quic: implement shutw Implement mux_ops shutw operation for QUIC mux. A RESET_STREAM is emitted unless the stream is already closed due to all data or RESET_STREAM already transmitted. This operation is notably useful when upper stream layer wants to close the connection early due to an error. This was tested by using a HTTP server which listens with PROXY protocol support. The corresponding server line on haproxy configuration deliberately not specify send-proxy. This causes the server to close abruptly the connection. Without this patch, nothing was done on the QUIC stream which was kept open until the whole connection is closed. Now, a proper RESET_STREAM is emitted to report the error. This should be backported up to 2.7.	2022-12-22 16:22:39 +01:00
William Lallemand	be6a873096	BUG/MINOR: httpclient/log: free of invalid ptr with httpclient_log_format free_proxy() must check if the ptr is not httpclient_log_format before trying to free p->conf.logformat_string. No backport needed.	2022-12-22 15:39:31 +01:00
William Lallemand	d793ca28b6	MEDIUM: httpclient: change the default log format to skip duplicate proxy data The httpclient emits logs in the httplog format, however it still display the frontend, the backend and the server. In the case of the httpclient we only need to know that we are using the httpclient, so the backend and server information are irelevant. In the case of extra code the name of the proxy can be long and will be displayed twice which is not useful. This is the same log-format as the httplog but the %b/%s is now -/- so the format is still compatible with an httplog parser. Before: <134>Dec 22 15:19:27 haproxy[1013520]: -:- [22/Dec/2022:15:19:27.482] <HTTPCLIENT> <HTTPCLIENT>/<HTTPCLIENT> 2/0/4/6/10 200 848 - - ---- 0/0/0/0/0 0/0 {92.123.236.161} "GET http://r3.o.lencr.org/1234 HTTP/1.1" After: <134>Dec 22 15:19:27 haproxy[1013520]: -:- [22/Dec/2022:15:19:27.482] <HTTPCLIENT> -/- 2/0/4/6/10 200 848 - - ---- 0/0/0/0/0 0/0 {92.123.236.161} "GET http://r3.o.lencr.org/1234 HTTP/1.1"	2022-12-22 15:13:59 +01:00
William Lallemand	a80b22eac4	MINOR: httpclient: don't add body when istlen is empty Don't try to create a request with a body in httpclient_req_gen() if the payload ist has a ptr but no len. Sometimes people have their httpclient stuck because they use an ist with a data ptr but no len. Check the len so this mistake doesn't block the client.	2022-12-22 14:49:43 +01:00
William Lallemand	70601c56da	BUG/MINOR: ssl/ocsp: httpclient blocked when doing a GET When the OCSP updater uses the GET method with the payload in the URI, the body must be set to IST_NULL, or the request won't be sent.	2022-12-22 14:41:31 +01:00
Christopher Faulet	c960a3b60f	BUG/MINOR: pool/stats: Use ullong to report total pool usage in bytes in stats The same change was already performed for the cli. The stats applet and the prometheus exporter are also concerned. Both use the stats API and rely on pool functions to get total pool usage in bytes. pool_total_allocated() and pool_total_used() must return 64 bits unsigned integer to avoid any wrapping around 4G. This may be backported to all versions.	2022-12-22 13:46:21 +01:00
Christopher Faulet	827a6299e6	BUG/MEDIUM: mux-h2: Refuse interim responses with end-stream flag set As state in RFC9113#8.1, HEADERS frame with the ES flag set that carries an informational status code is malformed. However, there is no test on this condition. On 2.4 and higher, it is hard to predict consequences of this bug because end of the message is only reported with a flag. But on 2.2 and lower, it leads to a crash because there is an unexpected extra EOM block at the end of an interim response. Now, when a ES flag is detected on a HEADERS frame for an interim message, a stream error is sent (RST_STREAM/PROTOCOL_ERROR). This patch should solve the issue #1972. It should be backported as far as 2.0.	2022-12-22 13:46:21 +01:00
William Lallemand	eb5302023f	CLEANUP: ssl/ocsp: add spaces around operators Add spaces around operators in ssl_ocsp_create_request_details().	2022-12-22 10:20:24 +01:00
William Lallemand	8bc00f8bdc	BUG/MINOR: ssl/ocsp: check chunk_strcpy() in ssl_ocsp_get_uri_from_cert() Check the return value of chunk_strcpy() in ssl_ocsp_get_uri_from_cert(). Should fix issue #1975.	2022-12-22 10:09:11 +01:00
Remi Tricot-Le Breton	c8d814ed63	MINOR: ssl: Move OCSP code to a dedicated source file This is a simple cleanup that moves OCSP related code to a dedicated file instead of interlacing it in some pure ssl connection code.	2022-12-21 11:21:07 +01:00
Remi Tricot-Le Breton	aff827785e	MEDIUM: ssl: Start update task if at least one ocsp-update option is set to on This patch effectively enables the ocsp auto update mechanism. If a least one ocsp-update option is enabled in a crt-list, then the ocsp auto update task is created. It will look into the dedicated ocsp update tree for the next update to be updated, use the http_client to send the ocsp request to the proper responder, validate the received ocsp response and update the ocsp response tree before finally reinserting the entry in the ocsp update tree (with a next update time set to now+1H). The main task will then sleep until another entry needs to be updated. The task gets scheduled after config check in order to avoid trying to update ocsp responses while configuration is still being parsed (and certificates and actual ocsp responses are loaded).	2022-12-21 11:21:07 +01:00
Remi Tricot-Le Breton	6477bbd78d	MEDIUM: ssl: Add ocsp update task main function This patch contains the main function of the ocsp auto update mechanism as well as an init and destroy function of the task used for this. The task is not created in this patch but in a later one. The function has two distinct parts and the branching to one or the other is completely based on the fact that the cur_ocsp pointer of the ssl_ocsp_task_ctx member is set. If the pointer is not set, we need to look at the first item of the update tree and see if it needs to be updated. If it does not we simply wait until the time is right and let the task asleep. If it does need to be updated, we simply build and send the corresponding ocsp request thanks to the http_client. The task is then sent to sleep with an expire time set to infinity. The http_client will wake it back up once the response is received (or a timeout occurs). Just note that during this whole process the cetificate_ocsp object corresponding to the entry being updated is taken out of the update tree and only stored in the ssl_ocsp_task_ctx context. Once the task is waken up by the http_client, it branches on the response processing part of the function which basically checks that the response is valid and inserts it into the ocsp_response tree. The task then goes back to sleep until another entry needs to be updated.	2022-12-21 11:21:07 +01:00
Remi Tricot-Le Breton	b55be8c90a	MEDIUM: ssl: Insert ocsp responses in update tree when needed When 'ocsp-update' is enabled for a given certificate, we need to insert the certificate_ocsp member of this certificate in the OCSP update tree as well as the already existing OCSP response tree. For such an entry to be created, the certificate needs to contain an "OCSP URI" field, and we also need to know the certificate's issuer, that is used to build the OCSP_CERTID. When no OCSP response is known for a given certificate, an empty certificate_ocsp object gets created so that it can be inserted in the ocsp update tree. The entry is inserted on the first spot of the update tree since its expire time is 0. Then whenever the update task is started, it will try to get responses for those certificates first. In order for the update process to work, we also need to store some information relative to the main certificate into the certificate_ocsp structure. This avoids having to keep a reference to a ckch in an ocsp tree entry. This patch adds a reference to the certificate chain as well as the ocsp issuer that might have been filled during init into the certificate_ocsp object. It also gets the ocsp uri at this time since it is contained in the server's certificate. We only take the first uri that might be contained in the certificate though. Those fields are only filled when ocsp auto update is enabled for the concerned certificate.	2022-12-21 11:21:07 +01:00
Remi Tricot-Le Breton	fb2b9988e8	MINOR: ssl: Store 'ocsp-update' mode in the ckch_data and check for inconsistencies The 'ocsp-update' option is parsed at the same time as all the other bind line options but it does not actually have anything to do with the bind line since it concerns the frontend certificate instead. For that reason, we should have a mean to identify inconsistencies in the configuration and raise an error when a given certificate has two different ocsp-update modes specified in one or more crt-lists. The simplest way to do it is to store the ocsp update mode directly in the ckch and not only in the ssl_bind_conf.	2022-12-21 11:21:07 +01:00
Remi Tricot-Le Breton	03c5ffff8e	MINOR: ssl: Add crt-list ocsp-update option This option will define how the ocsp update mechanism behaves. The option can either be set to 'on' or 'off' and can only be specified in a crt-list entry so that we ensure that it concerns a single certificate. The 'off' mode is the default one and corresponds to the old behavior (no automatic update). When the option is set to 'on', we will try to get an ocsp response whenever an ocsp uri can be found in the frontend's certificate. The only limitation of this mode is that the certificate's issuer will have to be known in order for the OCSP certid to be built. This patch only adds the parsing of the option. The full functionality will come in a later commit.	2022-12-21 11:21:07 +01:00
Remi Tricot-Le Breton	bdd3c79568	MINOR: ssl: Add ocsp_update_tree and helper functions The OCSP update tree holds ocsp responses that will need to be updated automatically. The entries are inserted in an eb64_tree where the keys are the absolute time after which the entry will need to be updated.	2022-12-21 11:21:07 +01:00
Remi Tricot-Le Breton	cc346678dc	MEDIUM: ssl: Add ocsp_certid in ckch structure and discard ocsp buffer early The ocsp_response member of the cert_key_and_chain structure is only used temporarily. During a standard init process where an ocsp response is provided, this ocsp file is first copied into the ocsp_response buffer without any ocsp-related parsing (see ssl_sock_load_ocsp_response_from_file), and then the contents are actually interpreted and inserted into the actual ocsp tree (cert_ocsp_tree) later in the process (see ssl_sock_load_ocsp). If the response was deemed valid, it is then copied into the actual ocsp_response structure's 'response' field (see ssl_sock_load_ocsp_response). From this point, the ocsp_response field of the cert_key_and_chain object could be discarded since actual ocsp operations will be based of the certificate_ocsp object. The only remaining runtime use of the ckch's ocsp_response field was in the CLI, and more precisely in the 'show ssl cert' mechanism. This constraint could be removed by adding an OCSP_CERTID directly in the ckch because the buffer was only used to get this id. This patch then adds the OCSP_CERTID pointer in the ckch, it clears the ocsp_response buffer early and simplifies the ckch_store_build_certid function.	2022-12-21 11:21:07 +01:00
Remi Tricot-Le Breton	eeaa29b36b	MINOR: ssl: Add "update ssl ocsp-response" cli command The new "update ssl ocsp-response <certfile>" CLI command allows to update the stored OCSP response for a given certificate. It relies on the http_client which is used to send an HTTP request to the OCSP responder whose URI can be extracted from the certificate. This command won't work for a certificate that did not have a stored OCSP response yet.	2022-12-21 11:21:07 +01:00
Remi Tricot-Le Breton	c0b4058e7e	MINOR: ssl: Add helper function that checks the validity of an OCSP response This helper function will check that an OCSP response is valid, meaning that the proper "Content-Type: application/ocsp-response" header is present and the data itself is a proper OCSP_RESPONSE that can be checked thanks to the issuer certificate.	2022-12-21 11:21:07 +01:00
Remi Tricot-Le Breton	e09d2ae598	MINOR: ssl: Add OCSP request helper function This function creates the url and body that will be used to build a proper OCSP request for a given certid (following section A.1 of RFC6960).	2022-12-21 11:21:07 +01:00
Remi Tricot-Le Breton	47a4f1239d	MINOR: ssl: Add helper function that extracts an OCSP URI from a certificate This function extracts the first OCSP URI (if any) contained in a certificate. It only takes the first of potentially multiple URIs.	2022-12-21 11:21:07 +01:00
Remi Tricot-Le Breton	95e7cf1ddf	MINOR: httpclient: Make the CLI flags public for future use Those flags used by the http_client in its CLI function might come to use for OCSP updates that will strongly rely on the http client.	2022-12-21 11:21:07 +01:00
Remi Tricot-Le Breton	2b96364b35	MINOR: ssl: Add a lock to the OCSP response tree The tree that contains OCSP responses is never locked despite being used at runtime for OCSP stapling as well as the CLI through "set ssl cert" and "set ssl ocsp-response" commands. Everything works though because the certificate_ocsp structure is refcounted and the tree's entries are cleaned up when SSL_CTXs are destroyed (thanks to an ex_data entry in which the certificate_ocsp pointer is stored). This new lock will come to use when the OCSP auto update mechanism is fully implemented because this new feature will be based on another tree that stores the same certificate_ocsp members and updates their contents periodically.	2022-12-21 11:21:07 +01:00
Willy Tarreau	8d49253588	BUG/MINOR: quic: do not allocate more rxbufs than necessary When QUIC thread binding was fixed by commit `f5a0c8abf` ("MEDIUM: quic: respect the threads assigned to a bind line"), one point was overlooked regarding rxbuf allocation. Indeed, there's one rxbuf per listener and per bound thread. Originally the loop would iterate over all threads, but this is not needed anymore and causes lots of memory to be allocated in scenarios where shards are used, the worst one being "shards by-thread" which allocates N^2 buffers for N threads. This gives us 2304 buffers (or 576 MB of RAM) for 48 threads! Let's only allocate one buffer per bound thread on each listener to fix this. This should be backported to 2.7 and generally wherever the commit above is backported. It depends on the previous commit below: "BUG/MEDIUM: quic: properly take shards into account on bind lines"	2022-12-21 09:27:26 +01:00
Willy Tarreau	eed7826529	BUG/MEDIUM: quic: properly take shards into account on bind lines Shards were completely forgotten in commit `f5a0c8abf` ("MEDIUM: quic: respect the threads assigned to a bind line"). The thread mask is taken from the bind_conf, but since shards were introduced in 2.5, the per-listener mask is held by the receiver and can be smaller than the bind_conf's mask. The effect here is that the traffic is not distributed to the appropriate thread. At first glance it's not dramatic since it remains one of the threads eligible by the bind_conf, but it still means that in some contexts such as "shards by-thread", some concurrency may persist on listeners while they're expected to be alone. One identified impact is that it requires more rxbufs than necessary, but there may possibly be other not yet identified side effects. This must be backported to 2.7 and everywhere the commit above is backported.	2022-12-21 09:27:26 +01:00
Amaury Denoyelle	15337fd808	BUG/MEDIUM: mux-quic: fix double delete from qcc.opening_list qcs instances for bidirectional streams are inserted in <qcc.opening_list>. It is removed from the list once a full HTTP request has been parsed. This is required to implement http-request timeout. In case a stream is deleted before receiving full HTTP request, it also must be removed from <qcc.opening_list>. This was not the case on first implementation but has been fixed by the following patch : `641a65ff3c` BUG/MINOR: mux-quic: remove qcs from opening-list on free This means that now a stream can be deleted from the list in two different functions. Sadly, as LIST_DELETE was used in both cases, nothing prevented a double-deletion from the list, even though LIST_INLIST was used. Both calls are replaced with LIST_DEL_INIT which is idempotent. This bug causes memory corruption which results in most cases in a segfault, most of times outside of mux-quic code itself. It has been found first by gabrieltz who reported it on the github issue #1903. Big thanks to him for his testing. This bug also causes failures on several 'M' transfer testcase of QUIC interop-runner. The s2n-quic client is particularly useful in this case as segfaults triggers were most of the times on the LIST_DELETE operation itself. This is probably due to its encapsulating of HEADERS frame with fin bit delayed in a following empty STREAM frame. This must be backported wherever the above patch is, up to 2.6.	2022-12-21 08:58:04 +01:00
Willy Tarreau	2aa14ce5a1	MINOR: pool: only use opportunistic versions of the swrate_add() functions We don't need to know very accurately how much RAM is needed in a pool, however we must not spend time competing with other threads trying to be the one with the most accurate value. Let's use the "_opportunistic" variants of swrate_add() which will simply cause some updates to be dropped in case of thread contention. This should significantly improve the situation when dealing with many threads and small per-thread caches. Performance gains of up to 1-2% were observed on 48-thread systems thanks to this alone.	2022-12-20 14:51:12 +01:00
Willy Tarreau	284cfc67b8	MINOR: pool: make the thread-local hot cache size configurable Till now it was only possible to change the thread local hot cache size at build time using CONFIG_HAP_POOL_CACHE_SIZE. But along benchmarks it was sometimes noticed a huge contention in the lower level memory allocators indicating that larger caches could be beneficial, especially on machines with large L2 CPUs. Given that the checks against this value was no longer on a hot path anymore, there was no reason for continuing to force it to be tuned at build time. So this patch allows to set it by tune.memory-hot-size. It's worth noting that during the boot phase the value remains zero so that it's possible to know if the value was set or not, which opens the possibility that we try to automatically adjust it based on the per-cpu L2 cache size or the use of certain protocols (none of this is done yet).	2022-12-20 14:51:12 +01:00
Christopher Faulet	a8b7684319	BUG/MEDIUM: stats: Rely on a local trash buffer to dump the stats It is possible to block the stats applet if a line exceeds the free space in the responsse buffer while the buffer is empty. It is only an issue in HTTP becaues of the HTX overhead and, AFAIK, only with json output. In this case, the applet is unable to write anything in the response buffer and waits for some free space to proceed further. On the other hand, because the response channel is empty, nothing is sent and thus no space can be freed. At this stage, the stream and the applet are blocked waiting for the other side. To avoid this situation, we must take care to not dump a line exceeding the free space in the HTX message. It means we cannot rely anymore on the global trash buffer. At least, not directly. The trick is to use a local trash buffer, mapped on the global one but with a different size. We use b_make() to do so. The local trash buffer is thread local to avoid any concurrency issue. It is a valid fix. However it could be good to review the internal API of the stats applet to not rely on a global variable. This patch should solve the #1873. It must be backported at least as far as 2.6. Older versions must be evaluated first but it is probably possible to hit this bug with long proxy/server names.	2022-12-19 11:01:26 +01:00
Christopher Faulet	ad4ed003f3	BUG/MINOR:: mux-h1: Never handle error at mux level for running connection During the request parsing, we must be sure to never handle errors at the mux level if the connection is running (or closing). The error must be handled by the upper layer. It should never happen. But there is an edge case that was not properly handled. If all data are received on the first packet with the read0 and the request is truncated on the payload, in this case the stream-connector is created, so the H1C is in RUNNING state. But an error is reported because the request is truncated. In this specific case, the error is handled by the mux while it should not. This patch is related to #1966. It must be backported to 2.7.	2022-12-19 11:01:26 +01:00
Christopher Faulet	75028f83d4	BUG/MINOR: mux-h1: Report EOS on parsing/internal error for not running stream When an error occurred during the request parsing while the stream is not running, an EOS must be reported. It is not an issue for an embryonic connection because the H1 stream is orphan. However, it is an issue with connections upgraded from TCP to H1. In this case, the upgrade is not performed because there is an early error. However the H1 stream is not orphan and is not destroyed. The H1 multiplexer will wait for the detach event. But without EOS, the upper layer is unable to perform the shutdown. This patch is related to #1966. It must be backported to 2.7. Older versions are not affected by this issue.	2022-12-19 11:01:26 +01:00
Amaury Denoyelle	5ac6b3b125	BUG/MINOR: quic: fix crash on PTO rearm if anti-amplification reset There is a possible segfault when accessing qc->timer_task in quic_conn_io_cb() without testing it. It seems however very rare as it requires several condition to be encounter. * quic_conn must be in CLOSING state after having sent a CONNECTION_CLOSE which free the qc.timer_task * quic_conn handshake must still be in progress : in fact, qc.timer_task is accessed on this path because of the anti-amplification limit lifted. I was unable thus far to trigger it but benchmarking tests seems to have fire it with the following backtrace as a result : #0 _task_wakeup (f=4096, caller=0x5620ed004a40 <_.46868>, t=0x0) at include/haproxy/task.h:195 195 state = _HA_ATOMIC_OR_FETCH(&t->state, f); [Current thread is 1 (Thread 0x7fc714ff1700 (LWP 14305))] (gdb) bt #0 _task_wakeup (f=4096, caller=0x5620ed004a40 <_.46868>, t=0x0) at include/haproxy/task.h:195 #1 quic_conn_io_cb (t=0x7fc5d0e07060, context=0x7fc5d0df49c0, state=<optimized out>) at src/quic_conn.c:4393 #2 0x00005620ecedab6e in run_tasks_from_lists (budgets=<optimized out>) at src/task.c:596 #3 0x00005620ecedb63c in process_runnable_tasks () at src/task.c:861 #4 0x00005620ecea971a in run_poll_loop () at src/haproxy.c:2913 #5 0x00005620ecea9cf9 in run_thread_poll_loop (data=<optimized out>) at src/haproxy.c:3102 #6 0x00007fc773c3f609 in start_thread () from /lib/x86_64-linux-gnu/libpthread.so.0 #7 0x00007fc77372d133 in clone () from /lib/x86_64-linux-gnu/libc.so.6 (gdb) up #1 quic_conn_io_cb (t=0x7fc5d0e07060, context=0x7fc5d0df49c0, state=<optimized out>) at src/quic_conn.c:4393 4393 task_wakeup(qc->timer_task, TASK_WOKEN_MSG); (gdb) p qc $1 = (struct quic_conn ) 0x7fc5d0df49c0 (gdb) p qc->timer_task $2 = (struct task ) 0x0 This fix should be backported up to 2.6.	2022-12-15 17:02:19 +01:00
Aurelien DARRAGON	16c9ca94ef	MINOR: stats: make show info json future-proof This is a follow up of "BUG/MINOR: stats: fix show stat json buffer limitation" However this time this is purely preemptive as we did not reach the buffer limitation yet. But now is the proper time so that this won't be an issue in the upcoming versions. No backport needed.	2022-12-15 16:53:49 +01:00
Aurelien DARRAGON	42b18fb645	BUG/MINOR: stats: fix show stat json buffer limitation json output type is a lot more verbose than other output types. Because of this and the increasing number of metrics implemented within haproxy, we are starting to reach max bufsize limit (defaults to 16k) when dumping stats to json since 2.6-dev1. This results in stats output being truncated with "[{"errorStr":"output buffer too short"}]" This was reported by Gabriel in #1964. Thanks to "MINOR: stats: introduce stats field ctx", we can now make multipart (using multiple buffers) dumping, in case a single buffer is not big enough to hold the complete stat line. For now, only stats_dump_fields_json() makes use of it as it is by far the most verbose stats output type. (csv, typed and html outputs should be good for a while and may use this capability if the need arises in some distant future) -- It could be backported to 2.6 and 2.7. This commit depends on: - MINOR: stats: provide ctx for dumping functions - MINOR: stats: introduce stats field ctx	2022-12-15 16:53:49 +01:00
Aurelien DARRAGON	5594184190	MINOR: stats: introduce stats field ctx Add a new value in stats ctx: field. Implement field support in line dumping parent functions stats_print_proxy_field_json() and stats_dump_proxy_to_buffer(). This will allow child dumping functions to support partial line dumping when needed. ie: when dumping buffer is exhausted: do a partial send and wait for a new buffer to finish the dump. Thanks to field ctx, the function can start dumping where it left off on previous (unterminated) invokation.	2022-12-15 16:53:49 +01:00
Aurelien DARRAGON	e76a027b0b	MINOR: stats: provide ctx for dumping functions This is a minor refactor to allow stats_dump_info_* and stats_dump_fields_* functions to directly access stat ctx pointer instead of explicitly passing stat ctx struct members to them. This will allow dumping functions to benefit from upcoming ctx updates.	2022-12-15 16:53:49 +01:00
Remi Tricot-Le Breton	4cf0d3f1e8	BUG/MINOR: ssl: Fix memory leak of find_chain in ssl_sock_load_cert_chain The certificate chain that gets passed in the SSL_CTX through SSL_CTX_set1_chain has its reference counter increased by OpenSSL itself. But since the ssl_sock_load_cert_chain function might create a brand new certificate chain if none exists in the ckch_data (sk_X509_new_null), then we ended up returning a new certificate chain to the caller that was never destroyed. This patch can be backported to all stable branches but it might need to be reworked for branches older than 2.4 because of commit `ec805a32b9` that refactorized the modified code.	2022-12-15 16:33:43 +01:00
Remi Tricot-Le Breton	e3d5f9a92b	MINOR: ssl: Remove unnecessary alloc'ed trash chunk in show ocsp-response When displaying the content of an OCSP response in the CLI, a buffer is alloc'ed when a temporary buffer would be enough.	2022-12-15 16:33:36 +01:00
Remi Tricot-Le Breton	9334843859	MINOR: ssl: Remove unneeded buffer allocation in show ocsp-response When calling 'show ssl ocsp-response' from the CLI, a temporary buffer was created in parse_binary when we could just use a local static buffer instead. This does not change the behavior of the function, it just simplifies it.	2022-12-15 16:33:25 +01:00
Amaury Denoyelle	c4913f6b54	MINOR: h3: check return values of htx_add_* on headers parsing Check return values of htx_add_header()/htx_add_eof() during H3 HEADERS conversion to HTX. In case of error, the connection is interrupted with a CONNECTION_CLOSE. This commit is useful to detect abnormal situation on headers parsing. It should be backported up to 2.7.	2022-12-15 11:48:30 +01:00
Amaury Denoyelle	788fc05401	BUG/MINOR: h3: fix memleak on HEADERS parsing failure If an error is triggered on H3 HEADERS parsing, the allocated buffer for HTX data is not freed. To prevent this memleak, all return path have been centralized using goto statements. Also, as a small bonus, offer_buffers() is not called anymore if buffer is not freed because sedesc has taken it. However this change has probably no noticeable effect as dynamic buffers management is not functional currently. This should be backported up to 2.6.	2022-12-15 11:48:30 +01:00
Amaury Denoyelle	19942e3859	BUG/MEDIUM: h3: fix cookie header parsing Cookie header are treated specifically to merge multiple occurences in a single HTX header. This is treated in a if-block condition inside the 'while (1)' loop for headers parsing. The length value of ist representing cookie header is set to -1 by http_cookie_register(). The problem is that then a continue statement is used but without incrementing 'hdr_idx' to pass on the next header. This issue was revealed by the introduction of commit : commit `d6fb7a0e0f` BUG/MEDIUM: h3: reject request with invalid header name Before the aformentionned patch, the bug was hidden : on the next while iteration, all isteq() invocations won't match with cookie header length now set to -1. htx_add_header() fails silently because length is invalid. hdr_idx is finally incremented which allows parsing to proceed normally with the next header. Now, a cookie header with length -1 do not pass the test on header name conformance introduced by the above patch. Thus, a spurrious RESET_STREAM is emitted. This behavior has been reported on the mailing list by Shawn Heisey who found out that browsers disabled H3 usage due to the RESET_STREAM received. Big thanks to him for his testing on the master branch. This issue is simply resolved by incrementing hdr_idx before continue statement. It could have been detected earlier if htx_add_header() return value was checked. This will be the subject of a dedicated commit outside of the backport scope. This must be backported up to 2.6.	2022-12-15 11:41:51 +01:00
Amaury Denoyelle	4328b61bb3	MINOR: http-htx: add BUG_ON to prevent API error on http_cookie_register http_cookie_register() must be called on first invocation with the last two arguments pointing both to a negative value. After it, they will be updated to a valid index. We must never have only the last argument as NULL as this will cause an invalid array addressing. To clarify this a BUG_ON statement is introduced. This is linked to github issue #1967.	2022-12-15 11:41:39 +01:00
Christopher Faulet	8f1f1b0579	BUG/MINOR: mux-h1: Fix test instead a BUG_ON() in h1_send_error() In the previous patch (86924532db "BUG/MINOR: mux-h1: Fix test instead a BUG_ON() in h1_send_error()"), a BUG_ON() condition was inverted by error in h1_send_error(). The stream-connector must be NULL to be able to destroy the H1 stream. This patch must be backported with the commit above (to 2.7).	2022-12-15 09:59:51 +01:00
Christopher Faulet	da93802ffc	BUG/MEDIUM: mux-h1: Don't release H1 stream upgraded from TCP on error When an error occurred during the request parsing, the H1 multiplexer is responsible to sent a response to the client and to release the H1 stream and the H1 connection. In HTTP mode, it is not an issue because at this stage the H1 connection is in embryonic state. Thus it can be released immediately. However, it is a problem if the connection was first upgraded from a TCP connection. In this case, a stream-connector is attached. The H1 stream is not orphan. Thus it must not be released at this stage. It must be detached first. Otherwise a BUG_ON() is triggered in h1s_destroy(). So now, the H1S is destroyed on early errors but only if the H1C is in embryonic state. This patch may be related to #1966. It must be backported to 2.7.	2022-12-15 09:51:31 +01:00
Amaury Denoyelle	d2c5ee665e	BUG/MEDIUM: h3: parse content-length and reject invalid messages Ensure that if a request contains a content-length header it matches with the total size of following DATA frames. This is conformance with HTTP/3 RFC 9114. For the moment, this kind of errors triggers a connection close. In the future, it should be handled only with a stream reset. To reduce backport surface, this will be implemented in another commit. This must be backported up to 2.6. It relies on the previous commit : MINOR: http: extract content-length parsing from H2	2022-12-14 11:37:12 +01:00
Amaury Denoyelle	15f3cc4b38	MINOR: http: extract content-length parsing from H2 Extract function h2_parse_cont_len_header() in the generic HTTP module. This allows to reuse it for all HTTP/x parsers. The function is now available as http_parse_cont_len_header(). Most notably, this will be reused in the next bugfix for the H3 parser. This is necessary to check that content-length header match the length of DATA frames. Thus, it must be backported to 2.6.	2022-12-14 11:34:18 +01:00
Amaury Denoyelle	7b5a671fb8	BUG/MEDIUM: h3: reject request with invalid pseudo header RFC 9114 dictates several requirements for pseudo header usage in H3 request. Previously only minimal checks were implemented. Enforce all the following requirements with this patch : * reject request with undefined or invalid pseudo header * reject request with duplicated pseudo header * reject non-CONNECT request with missing mandatory pseudo header * reject request with pseudo header after standard ones For the moment, this kind of errors triggers a connection close. In the future, it should be handled only with a stream reset. To reduce backport surface, this will be implemented in another commit. This must be backported up to 2.6.	2022-12-14 11:34:18 +01:00
Amaury Denoyelle	d6fb7a0e0f	BUG/MEDIUM: h3: reject request with invalid header name Reject request containing invalid header name. This concerns every header containing uppercase letter or a non HTTP token such as a space. For the moment, this kind of errors triggers a connection close. In the future, it should be handled only with a stream reset. To reduce backport surface, this will be implemented in another commit. Thanks to Yuki Mogi from FFRI Security, Inc. for having reported this. This must be backported up to 2.6.	2022-12-14 11:34:18 +01:00
Christopher Faulet	819d48b14e	BUG/MEDIUM: resolvers: Use tick_first() to update the resolvers task timeout In resolv_update_resolvers_timeout(), the resolvers task timeout is updated by checking running and waiting resolutions. However, to find the next wakeup date, MIN() operator is used to compare ticks. Ticks must never be compared with such operators, tick helper functions must be used, to properly handled TICK_ETERNITY value. In this case, tick_first() must be used instead of MIN(). It is an old bug but it is pretty visible since the commit `fdecaf6ae4` ("BUG/MINOR: resolvers: do not run the timeout task when there's no resolution"). Because of this bug, the resolvers task timeout may be set to TICK_ETERNITY, stopping periodic resolutions. This patch should solve the issue #1962. It must be backported to all stable versions.	2022-12-14 10:44:17 +01:00
Christopher Faulet	427293420f	BUG/MEDIUM: freq-ctr: Don't compute overshoot value for empty counters The function computing the excess of events of a frequency counter over the current period for a target frequency must handle empty counters (no event and no start date). In this case, no excess must be reported. Because of this bug, long pauses may be experienced on the bandwith limitation filter. This patch must be backported to 2.7.	2022-12-14 10:44:17 +01:00
William Lallemand	04007cb08d	CLEANUP: ssl: remove check on srv->proxy Remove a useless check on srv->proxy which triggers coverity. Should fix issue #1965.	2022-12-14 10:36:31 +01:00
Thayne McCombs	02cf4ecb5a	MINOR: sample: add param converter Add a converter that extracts a parameter from string of delimited key/value pairs. Fixes: #1697	2022-12-14 08:24:15 +01:00
William Lallemand	0adafb307e	BUG/MINOR: startup: don't use internal proxies to compute the maxconn With internal proxies using the SSL activated (httpclient for example) the automatic computation of the maxconn is wrong because these proxies are always activated by default. This patch fixes the issue by not counting these internal proxies during the computation. Must be backported as far as 2.5.	2022-12-13 18:28:29 +01:00
Amaury Denoyelle	4b167006fd	BUG/MINOR: mux-quic: handle properly alloc error in qcs_new() Use qcs_free() on allocation failure in qcs_new() This ensures that all qcs content is properly deallocated and prevent memleaks. Most notably, qcs instance is now removed from qcc tree. This bug is labelled as MINOR as it occurs only on qcs allocation failure due to memory exhaustion. This must be backported up to 2.6.	2022-12-12 14:54:39 +01:00
Amaury Denoyelle	641a65ff3c	BUG/MINOR: mux-quic: remove qcs from opening-list on free qcs instances for bidirectional streams are inserted in <qcc.opening_list>. It is removed from the list once a full HTTP request has been parsed. This is required to implement http-request timeout. If a qcs instance is freed before receiving a full HTTP request, it must be removed from the <qcc.opening_list>. Else a segfault will occur in qcc_refresh_timeout() when accessing a dangling pointer. For the moment this bug was not reproduced in production. This is because there exists only few rare cases where a qcs is freed before HTTP request parsing. However, as error detection will be improved on H3, this will occur more frequently in the near future. This must be backported up to 2.6.	2022-12-12 14:54:39 +01:00
Amaury Denoyelle	6eb3c4b71c	CLEANUP: mux-quic: remove unused attribute on qcs_is_close_remote() qcs_is_close_remote() is used in qcc_decode_qcs(). Thus the unused function attribute is now unneeded. This can be backported up to 2.7.	2022-12-12 14:54:39 +01:00
Amaury Denoyelle	4244833c5f	BUG/MINOR: quic: handle alloc failure on qc_new_conn() for owned socket This patch is the follow up of previous fix : BUG/MINOR: quic: properly handle alloc failure in qc_new_conn() quic_conn owned socket FD is initialized as soon as possible in qc_new_conn(). This guarantees that we can safely call quic_conn_release() on allocation failure. This function uses internally qc_release_fd() to free the socket FD unless it has been initialized to an invalid FD value. Without this patch, a segfault will occur if one inner allocation of qc_new_conn() fails before qc.fd is initialized. This change is linked to quic-conn owned socket implementation. This should be backported up to 2.7.	2022-12-12 14:53:55 +01:00
Amaury Denoyelle	dbf6ad470b	BUG/MINOR: quic: properly handle alloc failure in qc_new_conn() qc_new_conn() is used to allocate a quic_conn instance and its various internal members. If one allocation fails, quic_conn_release() is used to cleanup things. For the moment, pool_zalloc() is used which ensures that all content is null. However, some members must be initialized to a special values to be able to use quic_conn_release() safely. This is the case for quic_conn lists and its tasklet. Also, some quic_conn internal allocation functions were doing their own cleanup on failure without reset to NULL. This caused an issue with quic_conn_release() which also frees this members. To fix this, these functions now only return an error without cleanup. It is the caller responsibility to free the allocated content, which is done via quic_conn_release(). Without this patch, allocation failure in qc_new_conn() would often result in segfault. This was reproduced easily using fail-alloc at 10%. This should be backported up to 2.6.	2022-12-12 11:44:34 +01:00
Youfu Zhang	2e6bf0a272	BUG/MAJOR: fcgi: Fix uninitialized reserved bytes The output buffer is not zero-initialized. If we don't clear reserved bytes, fcgi requests sent to backend will leak sensitive data. This patch must be backported as far as 2.2.	2022-12-09 12:23:14 +01:00
Cedric Paillet	e06e31ea3b	MINOR: promex: introduce haproxy_backend_agg_check_status This patch introduces haproxy_backend_agg_check_status metric as we wanted in `42d7c402d` but with the right data source. This patch could be backported as far as 2.4.	2022-12-09 10:54:48 +01:00
Cedric Paillet	7d6644e689	BUG/MINOR: promex: create haproxy_backend_agg_server_status haproxy_backend_agg_server_check_status currently aggregates haproxy_server_status instead of haproxy_server_check_status. We deprecate this and create a new one, haproxy_backend_agg_server_status to clarify what it really does. This patch could be backported as far as 2.4.	2022-12-09 10:54:27 +01:00
Willy Tarreau	9192d20f02	MINOR: pools: make DEBUG_UAF a runtime setting Since the massive pools cleanup that happened in 2.6, the pools architecture was made quite more hierarchical and many alternate code blocks could be moved to runtime flags set by -dM. One of them had not been converted by then, DEBUG_UAF. It's not much more difficult actually, since it only acts on a pair of functions indirection on the slow path (OS-level allocator) and a default setting for the cache activation. This patch adds the "uaf" setting to the options permitted in -dM so that it now becomes possible to set or unset UAF at boot time without recompiling. This is particularly convenient, because every 3 months on average, developers ask a user to recompile haproxy with DEBUG_UAF to understand a bug. Now it will not be needed anymore, instead the user will only have to disable pools and enable uaf using -dMuaf. Note that -dMuaf only disables previously enabled pools, but it remains possible to re-enable caching by specifying the cache after, like -dMuaf,cache. A few tests with this mode show that it can be an interesting combination which catches significantly less UAF but will do so with much less overhead, so it might be compatible with some high-traffic deployments. The change is very small and isolated. It could be helpful to backport this at least to 2.7 once confirmed not to cause build issues on exotic systems, and even to 2.6 a bit later as this has proven to be useful over time, and could be even more if it did not require a rebuild. If a backport is desired, the following patches are needed as well: CLEANUP: pools: move the write before free to the uaf-only function CLEANUP: pool: only include pool-os from pool.c not pool.h REORG: pool: move all the OS specific code to pool-os.h CLEANUP: pools: get rid of CONFIG_HAP_POOLS DEBUG: pool: show a few examples in -dMhelp	2022-12-08 18:54:59 +01:00
Willy Tarreau	b634987fed	DEBUG: pool: show a few examples in -dMhelp It's not always easy to remember what certain options do together nor which ones are only relevant when combined with others, so let's add a few examples with the "help" command on -dM.	2022-12-08 18:45:41 +01:00
Willy Tarreau	4da51bd190	CLEANUP: pools: get rid of CONFIG_HAP_POOLS This one was set in defaults.h only when neither DEBUG_NO_POOLS nor DEBUG_UAF were set. This was not the most convenient location to look for it, and it was only used in pool.c to decide on the initial value of POOL_DBG_NO_CACHE. Let's just use DEBUG_NO_POOLS \|\| DEBUG_UAF directly on this flag and get rid of the intermediary condition. This also has the benefit of removing a double inversion, which is always nice for understanding.	2022-12-08 17:45:08 +01:00
Willy Tarreau	a95636682d	REORG: pool: move all the OS specific code to pool-os.h Till now pool-os used to contain a mapping from pool_{alloc,free}_area() to pool_{alloc,free}_area_uaf() in case of DEBUG_UAF, or the regular malloc-based function. And the _uaf() functions were in pool.c. But since 2.4 with the first cleanup of the pools, there has been no more calls to pool_{alloc,free}_area() from anywhere but pool.c, from exactly one place each. As such, there's no more need to keep _uaf() apart in pool.c, we can inline it into pool-os.h and leave all the OS stuff there, with pool.c calling either based on DEBUG_UAF. This is cleaner with less round trips between both files and easier to find.	2022-12-08 17:32:57 +01:00
Willy Tarreau	76a97a98ca	CLEANUP: pool: only include pool-os from pool.c not pool.h There's no need for the low-level pool functions to be known from all callers anymore, they're only used by pool.c. Let's reduce the amount of header files processed.	2022-12-08 17:32:40 +01:00
Willy Tarreau	67f89c527f	CLEANUP: pools: move the write before free to the uaf-only function In UAF mode, pool_put_to_os() performs a write to the about-to-be-freed memory area so as to make sure the page is properly mapped and catch a possible double-free. However there's no point keeping that in an ifdef in the generic function, because we now have a pool_free_area_uaf() that is the UAF-specific version of pool_free_area() and the one that is called immediately after this write. Let's move the code there, it will be cleaner.	2022-12-08 16:08:28 +01:00
William Lallemand	94dbfedec1	BUG/MEDIUM: httpclient/lua: double LIST_DELETE on end of lua task The lua httpclient cleanup can be called in 2 places, the hlua_httpclient_gc() and the hlua_httpclient_destroy_all(). A LIST_DELETE() is performed to remove the hlua_hc struct of the list. However, when the lua task ends and call hlua_ctx_destroy(), it does a LIST_DELETE() first, and then the gc tries to do a LIST_DELETE() again in hlua_httpclient_gc(), provoking a crash. This patch fixes the issue by doing a LIST_DEL_INIT() instead of LIST_DELETE() in both cases. Should fix issue #1958. Must be backported where `bb58142` is backported.	2022-12-08 11:30:03 +01:00
Willy Tarreau	57c3e75d4e	CLEANUP: init: remove useless assignment of nbthread The old test consisting in setting global.nbthread if lower than 1 is useless nowadays since it's already done in check_config_validity().	2022-12-08 08:14:35 +01:00
Willy Tarreau	400b3ae2d5	BUG/MINOR: init/threads: continue to limit default thread count to max per group Jakub Vojacek reported in issue #1955 that haproxy 2.7.0 doesn't start anymore on a 128-CPU machine with a default config. The reason is the raise of the default MAX_THREADS value that came with thread groups. Previously, the maximum number of threads was simply limited to this value, and all of them fit into one group. Now the limit being higher, all threads cannot fit by default into a single group, and haproxy fails to start. The solution adopted here is to continue to limit the number of threads to the max supported per group, but to multiply it by the number of groups (usually 1 by default). In addition, a diag warning is now emitted when this happens, reminding the user to set nbthread or adjust thread-groups. We can hardly do more than a diag warning if we don't want to make the upgrade painful for users. Thanks to Jakub for reporting this early. This must be backported to 2.7.	2022-12-08 08:14:35 +01:00
Aurelien DARRAGON	f648767a4e	MINOR: peers: unused code path in process_peer_sync In process_peer_sync: a check was performed to know whether the peers section handler should kill itself if the corresponding proxy was not started on the current process. This logic was initially implemented in early 1.6 development to prevent some issues when peers where used in conjunction with nbproc > 1: `f83d3fe00a` MEDIUM: init: stop any peers section not bound to the correct process `46dc1ca` MEDIUM: peers: unregister peers that were never started But later in 1.6 dev, a new commit has been introduced: `47c8c029db` MEDIUM: init: completely deallocate unused peers With the latter, the check implemented in `46dc1ca` ("MEDIUM: peers: unregister peers that were never started") will never succeed: it is dead code. Since nbproc support has been dropped in 2.5, things have changed a bit: `f83d3fe00a` logic was moved in mworker_cleanlisteners, but as in `46dc1ca` : peers task is safely destroyed before peers_fe is set to NULL. Conversely, peers_fe is first set by init_peers_frontend() before peers task is scheduled by peers_init_sync() in check_config_validity(). Again, it is safe to say that we will never reach !peers->peers_fe in process_peer_sync(): this self-killing mechanism is not relevant anymore. -- To cut a long story short: I stumbled on this while tracking down current signal api usage. This led me to a signal_unregister_handler() call performed in the aforementionned dead code. To me this code was potentially unsafe because signal_unregister_handler() is not thread safe and here it was used within a task initialized via task_new_anywhere(). So I decided to check how bad this could be (ie: conditions to be met for this code to run).. and here we are.	2022-12-07 18:26:53 +01:00
Aurelien DARRAGON	1412d31a6d	MINOR: mworker: remove unused legacy code in mworker_cleanlisteners This cleanup is a follow up of "CLEANUP: peers: unused code path in process_peer_sync" There are some remnants of 1.6 peers specific code in mworker_cleanlisteners() that was introduced with this patch serie: `f83d3fe00a` MEDIUM: init: stop any peers section not bound to the correct process `47c8c029db` MEDIUM: init: completely deallocate unused peers Back then, nbthread did not exist, nbproc was used instead. Updating some comments to make them more relevant to current haproxy design. (multithreaded single process) Moreover, in `47c8c029db`, task_free() was performed on peers_fe->task. But by looking at the code, from 1.6 til now, peers_fe->task is never used for peers proxies, it is only used for main proxies (referenced in proxies_list). Removing this extra task cleanup because it is misleading.	2022-12-07 18:26:53 +01:00
Aurelien DARRAGON	b118f2f407	MINOR: stats: properly handle ST_F_CHECK_DURATION metric ST_F_CHECK_DURATION metric is typed as unsigned int variable, and it is derived from check->duration that is signed. While most of the time check->duration > 0, it is not always true: with HCHK_STATUS_HANA checks, check->duration is set to -1 to prevent server logs from including irrelevant duration info (HCHK_STATUS_HANA checks are not time related). Because of this, stats could report UINT64_MAX value for ST_F_CHECK_DURATION metric. This was quite confusing. To prevent this, we make sure not to assign negative value to ST_F_CHECK_DURATION. This is only a minor printing issue, not backport needed.	2022-12-07 17:04:22 +01:00
Aurelien DARRAGON	81b7c9518c	MINOR: check: use atomic for s->consecutive_errors Properly use atomic operations when dealing with s->consecutive_errors as we're using it out of server's lock. Race is negligible, no backport needed.	2022-12-07 17:04:08 +01:00
Aurelien DARRAGON	7d541a91ec	BUG/MINOR: checks: restore legacy on-error fastinter behavior With previous commit, `9e080bf` ("BUG/MINOR: checks: make sure fastinter is used even on forced transitions"), on-error mark-down\|sudden-death\|fail-check are now working as expected. However, on-error fastinter remains broken because srv_getinter(), used in the above commit to check the expiration date, won't return fastinter interval if server health is maxed out (which is the case with on-error fastinter mode). To fix this, we introduce a check flag named CHK_ST_FASTINTER. This flag is set when on-error is triggered. This way we can force srv_getinter() to return fastinter interval whenever the flag is set. The flag is automatically cleared as soon as the new check task expiry is recalculated in process_chk_conn(). This restores original behavior prior to `d114f4a` ("MEDIUM: checks: spread the checks load over random threads"). It must be backported to 2.7 along with the aforementioned commits.	2022-12-07 17:03:55 +01:00
William Lallemand	e57b702e2b	BUG/MEDIUM: mworker: create the mcli_reload socketpairs in case of upgrade In ticket #1956, it was reported that an upgrade from 2.6 to 2.7 via a reload would stop the master process. When upgrading the binary, the new process is considered reexec and does not try to creates the socketpair for the mcli_reload listener, then tries to bind on -1 since the socket doesn't exit. The failure provokes an exit() of the master. This patch fixes the issue by trying to create the mcli_reload sockets only when they don't exist, instead of creating them at first start. This way we also avoid possible fd leak since we always try to use the existing FDs first. Must be backported in 2.7.	2022-12-07 15:30:52 +01:00
William Lallemand	035058e8bf	BUG/MEDIUM: mworker: fix segv in early failure of mworker mode with peers During an early failure of the mworker mode, the mworker_cleanlisteners() function is called and tries to cleanup the peers, however the peers are in a semi-initialized state and will use NULL pointers. The fix check the variable before trying to use them. Bug revealed in issue #1956. Could be backported as far as 2.0.	2022-12-07 15:27:36 +01:00
William Lallemand	40db4ae8bb	MINOR: mworker: display an alert upon a wait-mode exit When the mworker wait mode fails it does an exit, but there is no error message which says it exits. Add a message which specify that the error is non-recoverable. Could be backported in 2.7 and possibly earlier branch.	2022-12-07 15:07:53 +01:00
Ilya Shipitsin	5fa29b8a74	CLEANUP: assorted typo fixes in the code and comments This is 34th iteration of typo fixes	2022-12-07 09:08:18 +01:00
Willy Tarreau	9e080bf375	BUG/MINOR: checks: make sure fastinter is used even on forced transitions Aur�lien also found that while previous commit `a56798ea4` ("BUG/MEDIUM: checks: do not reschedule a possibly running task on state change") addressed one specific case where the check's task had to be woken up quickly, but it's not always sufficient as the check will not be considered as expired regarding the fastinter yet. Let's make sure we do consider this specific case to update the timer based on the new state if the new value is shorter. This particularly means that even if the timer is not expired yet during a wakeup when nothing is in progress, we need to check if applying the currently effective interval right now to the current date would expire earlier than what is programmed, then the timer needs to be updated. I.e. make sure we never miss fastinter during a state transition before the end of the current period. The approach is not pretty, but it forces to repass via the existing block dedicated to updating the timer if the current one is expired and the updated one would appear earlier. This must be backported to 2.7 along with the commit above.	2022-12-06 18:48:22 +01:00
Willy Tarreau	a56798ea4d	BUG/MEDIUM: checks: do not reschedule a possibly running task on state change Aur�lien found an issue introduced in 2.7-dev8 with commit `d114f4a68` ("MEDIUM: checks: spread the checks load over random threads"), but which in fact has deeper roots. When a server's state is changed via __health_adjust(), if a fastinter setting is set, the task gets rescheduled to run at the new date. The way it's done is not thread safe, as nothing prevents another thread where the task is already running from also updating the expire field in parallel. But since such events are quite rare, this statistically never happens. However, with the commit above, the tasks are no longer required to go to the shared wait queue and are no longer marked as shared between multiple threads. It's just that any thread may run them at a time without implying that all of them are allowed to modify them. And this change is sufficient to trigger the BUG_ON() condition in the scheduler that detects the inconsistency between a task queued in one thread and being manipulated in parallel by another one: FATAL: bug condition "task->tid != tid" matched at include/haproxy/task.h:670 call trace(13): \| 0x55f61cf520c9 [c6 04 25 01 00 00 00 00]: main-0x2ee7 \| 0x55f61d0646e8 [8b 45 08 a8 40 0f 85 65]: back_handle_st_cer+0x78/0x4d7 \| 0x55f61cff3e72 [41 0f b6 4f 01 e9 c8 df]: process_stream+0x2252/0x364f \| 0x55f61d0d2fab [48 89 c3 48 85 db 74 75]: run_tasks_from_lists+0x34b/0x8c4 \| 0x55f61d0d38ad [29 44 24 18 8b 54 24 18]: process_runnable_tasks+0x37d/0x6c6 \| 0x55f61d0a22fa [83 3d 0b 63 1e 00 01 0f]: run_poll_loop+0x13a/0x536 \| 0x55f61d0a28c9 [48 8b 1d f0 46 19 00 48]: main+0x14d919 \| 0x55f61cf56dfe [31 c0 e8 eb 93 1b 00 31]: main+0x1e4e/0x2d5d At first glance it looked like it could be addressed in the scheduler only, but in fact the problem clearly is at the application level, since some shared fields are manipulated without protection. At minima, the task's expiry ought to be touched only under the server's lock. While it's arguable that the scheduler could make such updates easier, changing it alone will not be sufficient here. Looking at the sequencing closer, it becomes obvious that we do not need this task_schedule() at all: a simple task_wakeup() is sufficient for the callee to update its timers. Indeed, the process_chk_con() function already deals with spurious wakeups, and already uses srv_getinter() to calculate the next wakeup date based on the current state. So here, instead of having to queue the task from __health_adjust() to anticipate a new check, we can simply wake the task up and let it decide when it needs to run next. This is much cleaner as the expiry calculation remains performed at a single place, from the task itself, as it should be, and it fixes the problem above. This should be backported to 2.7, but not to older versions where the risks of breakage are higher than the chance to fix something that ever happened.	2022-12-06 14:14:41 +01:00
Aurelien DARRAGON	22f82f81e5	MINOR: server/event_hdl: add support for SERVER_UP and SERVER_DOWN events We're using srv_update_status() as the only event source or UP/DOWN server events in an attempt to simplify the support for these 2 events. It seems srv_update_status() is the common path for server state changes anyway Tested with server state updated from various sources: - the cli - server-state file (maybe we could disable this or at least don't publish in global event queue in the future if it ends in slower startup for setups relying on huge server state files) - dns records (ie: srv template) (again, could be fined tuned to only publish in server specific subscriber list and no longer in global subscription list if mass dns update tend to slow down srv_update_status()) - normal checks and observe checks (HCHK_STATUS_HANA) (same as above, if checks related state update storms are expected) - lua scripts - html stats page (admin mode)	2022-12-06 10:22:07 +01:00
Aurelien DARRAGON	129ecf441f	MINOR: server/event_hdl: add support for SERVER_ADD and SERVER_DEL events Basic support for ADD and DEL server events are added through this commit: SERVER_ADD is published on dynamic server addition through cli. SERVER_DEL is published on dynamic server deletion through cli. This work depends on: "MINOR: event_hdl: add event handler base api" "MINOR: server: add srv->rid (revision id) value"	2022-12-06 10:22:07 +01:00
Aurelien DARRAGON	745ce8e8ad	MINOR: stats: add server revision id support Make use of the new srv->rid value in stats. Stat is referred as ST_F_SRID, it is now used in stats_fill_sv_stats function in order to be included in csv and json stats dumps. Moreover, "rid: $value" will be displayed next to server puid in html stats page if "stats show-legend" is specified in the stats frontend. (mouse hovering tooltip) Depends on the following commit: "MINOR: server: add srv->rid (revision id) value"	2022-12-06 10:22:06 +01:00
Aurelien DARRAGON	61e3894dfe	MINOR: server: add srv->rid (revision id) value With current design, we could not distinguish between previously existing deleted server and a new server reusing the deleted server name/id. This can cause some confusion when auditing stats/events/logs, because the new server will look similar to the old one. To address this, we're adding a new value in server structure: rid rid (revision id) value is an unsigned 32bits value that is set upon server creation. Value is derived from a global counter that starts at 0 and is incremented each time one or multiple server deletions are followed by a server addition (meaning that old name/id reuse could occur). Thanks to this revision id, it is now easy to tell whether the server we're looking at is the same as before or if it has been deleted and re-added in the meantime. (combining server name/id + server revision id yields a process-wide unique identifier)	2022-12-06 10:22:06 +01:00
Christopher Faulet	7f59d68fe2	BUG/MEDIIM: stconn: Flush output data before forwarding close to write side In process_stream(), we wait to have an empty output channel to forward a close to the write side (a shutw). However, at the stream-connector level, when a close is detected on one side and we don't want to keep half-close connections, the shutw is unconditionally forwarded to the write side. This typically happens on server side. At first glance, this bug may truncate messages. But depending on the muxes and the stream states, the bug may be more visible. On recent versions (2.8-dev and 2.7) and on 2.2 and 2.0, the stream may be freezed, waiting for the client timeout, if the client mux is unable to forward data because the client is too slow _AND_ the response channel is not empty _AND_ the server closes its connection _AND_ the server mux has forwarded all data to the upper layer _AND_ the client decides to send some data and to close its connection. On 2.6 and 2.4, it is worst. Instead of a freeze, the client mux is woken up in loop. Of course, conditions are pretty hard to meet. Especially because it is highly time dependent. For what it's worth, I reproduce it with tcploop on client and server sides and a basic HTTP configuration for HAProxy: * client: tcploop -v 8889 C S:"GET / HTTP/1.1\r\nConnection: upgrade\r\n\r\n" P5000 S:"1234567890" K * server: tcploop -v 8000 L A R S:"HTTP/1.1 101 ok\r\nConnection: upgrade\r\n\r\n" P2000 S2660000 F R On 2.8-dev, without this patch, the stream is freezed and when the client connection timed out, client data are truncated and '--cL' is reported in logs. With the patch, the client data are forwarded to the server and the connection is closed. A '--CD' is reported in logs. It is an old bug. It was probably introduced with the multiplexers. To fix it, in stconn (Formerly the stream-interface), we must wait all output data be flushed before forwarding close to write side. This patch must be backported as far as 2.2 and must be evaluated for 2.0.	2022-12-05 11:24:24 +01:00
Amaury Denoyelle	30fc27750d	BUG/MINOR: quic: fix fd leak on startup check quic-conn owned socket A startup check is done for first QUIC listener to detect if quic-conn owned socket is supported by the system. This is done by creating a dummy socket reusing the listener address. This socket must be closed as soon as the check is done. The socket condition is invalid as it excludes zero which is a valid file-descriptor value. Fix this bug by adjusting this condition. In theory, this bug could prevent the usage of quic-conn owned socket as startup check would report a false error. Also, the file-descriptor would leak as it is not closed. In practice, this cannot happen when startup check is done after a 'quic4/quic6' listener is instantiated as file-descriptor are allocated in ascending order by the system. This should fix github issue #1954. quic-conn owned socket implementation is scheduled for backport on 2.7. This commit must be backported with it, more specifically to fix the following patch : `75839a44e7` MINOR: quic: startup detect for quic-conn owned socket support	2022-12-05 10:45:20 +01:00
William Lallemand	151dbbe778	BUG/MINOR: ssl: initialize WolfSSL before parsing The wolfSSL library need to be initialized before parsing the configuration which uses some SSL functions. To be backported in 2.6.	2022-12-02 17:17:43 +01:00
William Lallemand	44c80ce5b3	BUG/MINOR: ssl: initialize SSL error before parsing The SSL error initialization need to be done before the configuration parsing, because it uses the SSL. Need to be backported to 2.6.	2022-12-02 17:10:11 +01:00
Amaury Denoyelle	e30f378236	MINOR: quic: activate socket per conn by default Activate QUIC connection socket to achieve the best performance. The previous behavior can be reverted by tune.quic.socket-owner configuration option. This change is part of quic-conn owned socket implementation. Contrary to its siblings patches, I suggest to not backport it to 2.7. This should ensure that stable releases behavior is perserved. If a user faces issues with QUIC performance on 2.7, he can nonetheless change the default configuration.	2022-12-02 14:45:43 +01:00
Amaury Denoyelle	d3083c9df9	MINOR: quic: reconnect quic-conn socket on address migration UDP addresses may change over time for a QUIC connection. When using quic-conn owned socket, we have to detect address change to break the bind/connect association on the socket. For the moment, on change detected, QUIC connection socket is closed and a new one is opened. In the future, we may improve this by trying to keep the original socket and reexecute only bind/connect syscalls. This change is part of quic-conn owned socket implementation. It may be backported to 2.7 after a period of observation.	2022-12-02 14:45:43 +01:00
Amaury Denoyelle	b2bd83972b	MEDIUM: quic: requeue datagrams received on wrong socket There is a small race condition when QUIC connection socket is instantiated between the bind() and connect() system calls. This means that the first datagram read on the sockets may belong to another connection. To detect this rare case, we compare the DCID for each QUIC datagram read on the QUIC socket. If it does not match the connection CID, the datagram is requeue using quic_receiver_buf to be able to handle it on the correct thread. This change is part of quic-conn owned socket implementation. It may be backported to 2.7 after a period of observation.	2022-12-02 14:45:43 +01:00
Amaury Denoyelle	b7ce79c814	MINOR: mux-quic: rename duplicate function names qc_rcv_buf and qc_snd_buf are names used for static functions in both quic-sock and quic-mux. To remove this ambiguity, slightly modify names used in MUX code. In the future, we should properly define a unique prefix for all QUIC MUX functions to avoid such problem in the future. This change is part of quic-conn owned socket implementation. It may be backported to 2.7 after a period of observation.	2022-12-02 14:45:43 +01:00
Amaury Denoyelle	7c9fdd9c3a	MEDIUM: quic: move receive out of FD handler to quic-conn io-cb This change is the second part for reception on QUIC connection socket. All operations inside the FD handler has been delayed to quic-conn tasklet via the new function qc_rcv_buf(). With this change, buffer management on reception has been simplified. It is now possible to use a local buffer inside qc_rcv_buf() instead of quic_receiver_buf(). This change is part of quic-conn owned socket implementation. It may be backported to 2.7 after a period of observation.	2022-12-02 14:45:43 +01:00
Amaury Denoyelle	5b41486b7f	MEDIUM: quic: use quic-conn socket for reception Try to use the quic-conn socket for reception if it is allocated. For this, the socket is inserted in the fdtab. This will call the new handler quic_conn_io_cb() which is responsible to process the recv() system call. It will reuse datagram dispatch for simplicity. However, this is guaranteed to be called on the quic-conn thread, so it will be more efficient to use a dedicated buffer. This will be implemented in another commit. This patch should improve performance by reducing contention on the receiver socket. However, more gain can be obtained when the datagram dispatch operation will be skipped. Older quic_sock_fd_iocb() is renamed to quic_lstnr_sock_fd_iocb() to emphasize its usage for the receiver socket. This change is part of quic-conn owned socket implementation. It may be backported to 2.7 after a period of observation.	2022-12-02 14:45:43 +01:00
Amaury Denoyelle	dc0dcb394b	MINOR: quic: use connection socket for emission If quic-conn has a dedicated socket, use it for sending over the listener socket. This should improve performance by reducing contention over the shared listener socket. This change is part of quic-conn owned socket implementation. It may be backported to 2.7 after a period of observation.	2022-12-02 14:45:43 +01:00
Amaury Denoyelle	40909dfec5	MINOR: quic: allocate a socket per quic-conn Allocate quic-conn owned socket if possible. This requires that this is activated in haproxy configuration. Also, this is done only if local address is known so it depends on the support of IP_PKTINFO. For the moment this socket is not used. This causes QUIC support to be broken as received datagram are not read. This commit will be completed by a following patch to support recv operation on the newly allocated socket. This change is part of quic-conn owned socket implementation. It may be backported to 2.7 after a period of observation.	2022-12-02 14:45:43 +01:00
Amaury Denoyelle	511ddd5785	MINOR: quic: define config option for socket per conn Define global configuration option "tune.quic.socket-owner". This option can be used to activate or not socket per QUIC connection mode. The default value is "listener" which disable this feature. It can be activated with the option "connection". This change is part of quic-conn owned socket implementation. It may be backported to 2.7 after a period of observation.	2022-12-02 14:45:43 +01:00
Amaury Denoyelle	8d46acdfcb	MINOR: quic: test IP_PKTINFO support for quic-conn owned socket Extend the startup platform detection support test for quic-conn owned socket. It is required to be able to retrieve destination address on a recvfrom() system call so check if IP_PKTINFO or IP_RECVDSTADDR flags are supported. This change is part of quic-conn owned socket implementation. It may be backported to 2.7 after a period of observation.	2022-12-02 14:45:43 +01:00
Amaury Denoyelle	75839a44e7	MINOR: quic: startup detect for quic-conn owned socket support To be able to use individual sockets for QUIC connections, we rely on the OS network stack which must support UDP sockets binding on the same local address. Add a detection code for this feature executed on startup. When the first QUIC listener socket is binded, a test socket is created and binded on the same address. If the bind call fails, we consider that it's impossible to use individual socket for QUIC connections. A new global option GTUNE_QUIC_SOCK_PER_CONN is defined. If startup detect fails, this value is resetted from global options. For the moment, there is no code to activate the option : this will be in a follow-up patch with the introduction of a new configuration option. This change is part of quic-conn owned socket implementation. It may be backported to 2.7 after a period of observation.	2022-12-02 14:45:43 +01:00
Amaury Denoyelle	eb6be98a65	MINOR: quic: ignore address migration during handshake QUIC protocol support address migration which allows to maintain the connection even if client has changed its network address. This is done through address migration. RFC 9000 stipulates that address migration is forbidden before handshake has been completed. Add a check for this : drop silently every datagram if client network address has changed until handshake completion. This commit is one of the first steps towards QUIC connection migration support. This should be backported up to 2.7.	2022-12-02 14:45:43 +01:00
Amaury Denoyelle	eec0b3c1bd	MINOR: quic: detect connection migration Detect connection migration attempted by the client. This is done by comparing addresses stored in quic-conn with src/dest addresses of the UDP datagram. A new function qc_handle_conn_migration() has been added. For the moment, no operation is conducted and the function will be completed during connection migration implementation. The only notable things is the increment of a new counter "quic_conn_migration_done". This should be backported up to 2.7.	2022-12-02 14:45:43 +01:00
Amaury Denoyelle	21e611dc89	MINOR: tools: add port for ipcmp as optional criteria Complete ipcmp() function with a new argument <check_port>. If this argument is true, the function will compare port values besides IP addresses and return true only if both are identical. This commit will simplify QUIC connection migration detection. As such, it should be backported to 2.7.	2022-12-02 14:45:43 +01:00
Amaury Denoyelle	8687b63c69	MINOR: quic: extract datagram parsing code Extract individual datagram parsing code outside of datagrams list loop in quic_lstnr_dghdlr(). This is moved in a new function named quic_dgram_parse(). To complete this change, quic_lstnr_dghdlr() has been moved into quic_sock source file : it belongs to QUIC socket lower layer and is directly called by quic_sock_fd_iocb(). This commit will ease implementation of quic-conn owned socket. New function quic_dgram_parse() will be easily usable after a receive operation done on quic-conn IO-cb. This should be backported up to 2.7.	2022-12-02 14:45:43 +01:00
Amaury Denoyelle	3f474e64c8	MINOR: quic: complete traces in qc_rx_pkt_handle() Add missing ENTER trace for qc_rx_pkt_handle() function. LEAVE traces are already present. This should be backported up to 2.7.	2022-12-02 14:45:43 +01:00
Amaury Denoyelle	518c98f150	MINOR: quic: remove qc from quic_rx_packet quic_rx_packet struct had a reference to the quic_conn instance. This is useless as qc instance is always passed through function argument. In fact, pkt.qc is used only in qc_pkt_decrypt() on key update, even though qc is also passed as argument. Simplify this by removing qc field from quic_rx_packet structure definition. Also clean up qc_pkt_decrypt() documentation and interface to align it with other quic-conn related functions. This should be backported up to 2.7.	2022-12-02 14:45:43 +01:00
William Lallemand	52ddd99940	MEDIUM: ssl: rename the struct "cert_key_and_chain" to "ckch_data" Rename the structure "cert_key_and_chain" to "ckch_data" in order to avoid confusion with the store whcih often called "ckchs". The "cert_key_and_chain ckch" were renamed "ckch_data data", so we now have store->data instead of ckchs->ckch. Marked medium because it changes the API.	2022-12-02 11:48:30 +01:00
Aurelien DARRAGON	68e692da02	MINOR: event_hdl: add event handler base api Adding base code to provide subscribe/publish API for internal events processing. event_hdl provides two complementary APIs, both are implemented in src/event_hdl.c and include/haproxy/event_hdl{-t.h,.h}: One API targeting developers that want to register event handlers that will be notified on specific events. (SUBSCRIBE) One API targeting developers that want to notify registered handlers about an event. (PUBLISH) This feature is being considered to address the following scenarios: - mailers code refactoring (getting rid of deprecated tcp-check ruleset implementation) - server events from lua code (registering user defined lua function that is executed with relevant data when a server is dynamically added/removed or on server state change) - providing a stable and easy to use API for upcoming developments that rely on specific events to perform actions. (e.g: ressource cleanup when a server is deleted from haproxy) At this time though, we don't have much use cases in mind in addition to server events handling, but the API is aimed at being multipurpose so that new event families, with their own particularities, can be easily implemented afterwards (and hopefully) without requiring breaking changes to the API. Moreover, you should know that the API was not designed to cope well with high rate event publishing. Mostly because publishing means iterating over unsorted subscriber list. So it won't scale well as subscriber list increases, but it is intended in order to keep the code simple and versatile. Instead, it is assumed that events implemented using this API should be periodic events, and that events related to critical io/networking processing should be handled using dedicated facilities anyway. (After all, this is meant to be a general purpose event API) Apart from being easily extensible, one of the main goals of this API is to make subscriber code as simple and safe as possible. This is done by offering multiple event handling modes: - SYNC mode: publishing code directly leverages handler code (callback function) and handler code has a direct access to "live" event data (pointers mostly, alongside with lock hints/context so that accessing data pointers can be done properly) - normal ASYNC mode: handler is executed in a backward compatible way with sync mode, so that it is easy to switch from and to SYNC/ASYNC mode. Only here the handler has access to "offline" event data, and not "live" data (ptrs) so that data consistency is guaranteed. By offline, you should understand "snapshot" of relevant data at the time of the event, so that the handler can consume it later (even if associated ressource is not valid anymore) - advanced ASYNC mode same as normal ASYNC mode, but here handler is not a function that is executed with event data passed as argument: handler is a user defined tasklet that is notified when event occurs. The tasklet may consume pending events and associated data through its own message queue. ASYNC mode should be considered first if you don't rely on live event data and you wan't to make sure that your code has the lowest impact possible on publisher code. (ie: you don't want to break stuff) Internal API documentation will follow: You will find more details about the notions we roughly approached here.	2022-12-02 09:40:52 +01:00
Willy Tarreau	b59e3f6045	MINOR: debug: add a balance of alloc - free at the end of the memstats dump When digging into suspected memory leaks, it's cumbersome to count the number of allocations and free calls. Here we're adding a summary at the end of the sum of allocs minus the sum of frees, excluding realloc since we can't know how much it releases upon each call. This means that when doing many realloc+free the count may be negative but in practice there are very few reallocs so that's not a problem. Also the size/call is signed and corresponds to the average size allocated (e.g. leaked) per call. It seems to work reasonably well for now: > debug dev memstats match buf quic_conn.c:2978 P_FREE size: 1239547904 calls: 75656 size/call: 16384 buffer quic_conn.c:2960 P_ALLOC size: 1239547904 calls: 75656 size/call: 16384 buffer mux_quic.c:393 P_ALLOC size: 9112780800 calls: 556200 size/call: 16384 buffer mux_quic.c:383 P_ALLOC size: 17783193600 calls: 1085400 size/call: 16384 buffer mux_quic.c:159 P_FREE size: 8935833600 calls: 545400 size/call: 16384 buffer mux_quic.c:142 P_FREE size: 9112780800 calls: 556200 size/call: 16384 buffer h3.c:776 P_ALLOC size: 8935833600 calls: 545400 size/call: 16384 buffer quic_stream.c:166 P_FREE size: 975241216 calls: 59524 size/call: 16384 buffer quic_stream.c:127 P_FREE size: 7960592384 calls: 485876 size/call: 16384 buffer stream.c:772 P_FREE size: 8798208 calls: 537 size/call: 16384 buffer stream.c:768 P_FREE size: 2424832 calls: 148 size/call: 16384 buffer stream.c:751 P_ALLOC size: 8852062208 calls: 540287 size/call: 16384 buffer stream.c:641 P_FREE size: 8849162240 calls: 540110 size/call: 16384 buffer stream.c:640 P_FREE size: 8847360000 calls: 540000 size/call: 16384 buffer channel.h:850 P_ALLOC size: 2441216 calls: 149 size/call: 16384 buffer channel.h:850 P_ALLOC size: 5914624 calls: 361 size/call: 16384 buffer dynbuf.c:55 P_FREE size: 32768 calls: 2 size/call: 16384 buffer Total BALANCE size: 0 calls: 5606906 size/call: 0 (excl. realloc) Let's see how useful this becomes over time.	2022-12-01 16:12:21 +01:00
Willy Tarreau	e57fbed3c4	MINOR: debug: support pool filtering on "debug dev memstats" Sometimes when debugging it's convenient to be able to focus only on certain pools. Just like we did for "show pools", let's add a filter based on a prefix on "debug dev memstats match <prefix>".	2022-12-01 16:12:21 +01:00
Willy Tarreau	111c78329e	MINOR: debug: relax access restrictions on "debug dev hash" and "memstats" These two have absolutely zero impact on the process and do not need to be restricted to the expert mode. The first one calculates a string hash that can be used by anyone when checking a dump; the second one may be used by anyone tracking a memory leak, and is cumbersome to use due to the "expert-mode on" that needs to be prepended. In addition this gives bad habits to users and needlessly taints the process. So let's drop this restriction for these two commands.	2022-11-30 17:58:00 +01:00
Willy Tarreau	50dd7e95c8	CLEANUP: anon: clarify the help message on "debug dev hash" This command is used to hash a section name using the current anon key, it was brought in 2.7 by commit `54966dffd` ("MINOR: anon: store the anonymizing key in the CLI's appctx"). However the help message only says "return msg hashed" which is misleading because if anon mode is not enabled, it returns the string as-is. Let's just mention this condition in the help message, and also fix the alphabetical ordering and alignment on the line.	2022-11-30 17:58:00 +01:00
Willy Tarreau	334d091b75	MINOR: debug: improve error handling on the memstats command parser "debug dev memstats" supports various options but silently ignores the unknown ones. Let's make sure it returns indications about what it expects, as the help message is quite limited otherwise.	2022-11-30 17:24:29 +01:00
Christopher Faulet	38f61351c3	MINOR: mux-h1: add the expire task and its expiration date in "show fd" Just like for the H2 multiplexer, info about the H1 connection task is now displayed in "show fd" output. The task pointer is displayed and, if not null, its expiration date. It may be useful to backport it.	2022-11-30 14:49:56 +01:00
Ilya Shipitsin	6f86eaae4f	CLEANUP: assorted typo fixes in the code and comments This is 33rd iteration of typo fixes	2022-11-30 14:02:36 +01:00
Willy Tarreau	4ede46be4e	BUG/MINOR: peers: always update the stksess shard number on incoming updates If shards are in use, we must fill the shard number on incoming updates, otherwise some entries are assigned shard number zero, and may be broadcast everywhere once updated, instead of being sent only to the peers having the same shard number. This fixes commit `36d156564` ("MINOR: peers: Support for peer shards"). No backport is needed.	2022-11-29 18:06:42 +01:00
Willy Tarreau	b12be7c1bb	CLEANUP: peers: factor out the key len calculation in received updates In peer_treat_updatemsg(), the lower layers of the stick-table code are reimplemented, and the key length is never really known for an entry being processed, it depends on the type being parsed and the moment where it's done. This makes it quite difficult to stuff some shard number calculation there. This patch adds a keylen local variable that is always set to the length of the current key depending on its type. It takes this opportunity for reducing redudant expressions involving this length and always using the new variable instead, limiting the risk of errors. Arguably that code would have been way simpler by creating a dummy stktable_key and passing it to stksess_new() as done anywhere else, but let's not change all that a few days before the release.	2022-11-29 18:06:42 +01:00
Willy Tarreau	d5cae6a0c7	MINOR: stick-table: change the API of the function used to calculate the shard The function used to calculate the shard number currently requires a stktable_key on input for this. Unfortunately, it happens that peers currently miss this calculation and they do not provide stktable_key at all, instead they're open-coding all the low-level stick-table work (hence why it's missing). Thus we'll need to be able to calculate the shard number in keys coming from peers as well but the current API does not make it possible. This commit addresses this by inverting the order where the length and the shard number are used. Now the low-level function is independent on stksess and stktable_key, it takes a table, pointer and length and does all the job. The upper function takes care of the type and key to get the its length, and is for use only from stick-table code. This doesn't change anything except that the low-level one will be usable from outside (hence why it's exported now).	2022-11-29 18:06:42 +01:00
Christopher Faulet	061a098c5c	BUG/MEDIUM: mux-h1: Close client H1C on EOS when there is no output data If the client closes the connection while there is no pending outgoing data, the H1 connection must be released. However, it was switched to CLOSING state instead. Thus the client connection was closed on client timeout. It is side effect of the commif `d1b573059a` ("MINOR: mux-h1: Avoid useless call to h1_send() if no error is sent"). Before, the extra call to h1_send() was able to fix the H1C state. To fix the bug and make switch to close state (CLOSING or CLOSED) less errorprone, h1_close() helper function is systematically used. It is a 2.7-specific bug. No backport needed.	2022-11-29 17:25:02 +01:00
Willy Tarreau	e548a7af45	BUG/MINOR: peers: always initialize the stksess shard value We need to initialize the shard value in __stksess_init() because there is not necessarily a key to make it happen later, resulting in an uninitialized shard value appearing in the entry, typically when entries are learned from peers. This fixes commit `36d156564` ("MINOR: peers: Support for peer shards"), no backport is needed. Note however that it is not sufficient to completely fix the peers code, the shard value remains zero because the setting of the key was open-coded in the peers code and these parts were not identified when adding support for shards.	2022-11-29 16:33:37 +01:00
Willy Tarreau	f8c7709013	MINOR: mux-h2: add the expire task and its expiration date in "show fd" Some issues such as #1929 seem to involve a task without timeout but we can't find the condition to reproduce this in the code. However, not having this info in the output doesn't help, so this patch adds the task pointer and its timeout (when the task is non-null). It may be useful to backport it.	2022-11-29 15:29:00 +01:00
Frédéric Lécaille	7b5d9b1f03	BUG/MINOR: quic: Endless loop during retransmissions qc_dgrams_retransmit() could reuse the same local list and could splice it two times to the packet number space list of frame to be send/resend. This creates a loop in this list and makes qc_build_frms() possibly endlessly loop when trying to build frames from the packet number space list of frames. Then haproxy aborts. This issue could be easily reproduced patching qc_build_frms() function to set <dlen> variable value to 0 after having built at least 10 CRYPTO frames and using ngtcp2 as client with 30% packet loss in both direction. Thank you to @gabrieltz for having reported this issue in GH #1903. Must be backported to 2.6.	2022-11-29 15:19:16 +01:00
Amaury Denoyelle	2526a6aca5	CLEANUP: ncbuf: use standard BUG_ON with DEBUG_STRICT ncbuf can be compiled for haproxy or standalone to run unit test suite. For the latest mode, BUG_ON() macro has been re-implemented in a simple version. The inclusion of the default or the redefined macro relied on DEBUG_DEV. Change this to now rely on DEBUG_STRICT as this is activated for the default build. This change is safe as only BUG_ON_HOT() macro is used in ncbuf code, which is activated only with the default value DEBUG_STRICT=2. This should be backported up to 2.6.	2022-11-29 15:15:27 +01:00
Amaury Denoyelle	d64a26f023	CLEANUP: ncbuf: inline small functions ncbuf API relies on lot of small functions. Mark these functions as inline to reduce call invocations and facilitate compiler optimizations to reduce code size. This should be backported up to 2.6.	2022-11-29 15:14:39 +01:00
Amaury Denoyelle	17e20e8cef	CLEANUP: ncbuf: remove ncb_blk args by value ncb_blk structure is used to represent a block of data or a gap in a non-contiguous buffer. This is used in several functions for ncbuf implementation. Before this patch, ncb_blk was passed by value, which is sub-optimal. Replace this by const pointer arguments. This has the side-effect of suppressing a compiler warning reported in older GCC version : CC src/http_conv.o src/ncbuf.c: In function 'ncb_blk_next': src/ncbuf.c:170: warning: 'blk.end' may be used uninitialized in this function This should be backported up to 2.6.	2022-11-29 15:12:54 +01:00
Willy Tarreau	16b282f4b0	MINOR: stick-table: show the shard number in each entry's "show table" output Stick-tables support sharding to multiple peers but there was no way to know to what shard an entry was going to be sent. Let's display this in the "show table" output to ease debugging.	2022-11-29 12:00:49 +01:00
Willy Tarreau	56460ee52a	MINOR: stick-table: store a per-table hash seed and use it Instead of using memcpy() to concatenate the table's name to the key when allocating an stksess, let's compute once for all a per-table seed at boot time and use it to calculate the key's hash. This saves two memcpy() and the usage of a chunk, it's always nice in a fast path. When tested under extreme conditions with a 80-byte long table name, it showed a 1% performance increase.	2022-11-28 18:58:06 +01:00
Willy Tarreau	f9607f8b1f	REORG: activity/cli: move the "show activity" handler to activity.c Initially the code was placed into cli.c to keep activity.c small and independent of the cli stuff, but that's no longer the case anyway and keeping that code over there makes it harder to find. Let's move it to its more natural place now.	2022-11-25 15:41:47 +01:00
Willy Tarreau	04b5b266e5	MINOR: activity: report uptime in "show activity" It happened a few times that it was difficult to figure if a counter was normal or not in "show activity" based on the uptime. Let's just emit the uptime value along with the date.	2022-11-25 15:36:48 +01:00
William Lallemand	0a2d63236c	BUG/MINOR: ssl: shut the ca-file errors emitted during httpclient init With an OpenSSL library which use the wrong OPENSSLDIR, HAProxy tries to load the OPENSSLDIR/certs/ into @system-ca, but emits a warning when it can't. This patch fixes the issue by allowing to shut the error when the SSL configuration for the httpclient is not explicit. Must be backported in 2.6.	2022-11-24 19:14:19 +01:00
William Lallemand	3992f55ff3	MINOR: ssl: forgotten newline in error messages on ca-file Add forgotten newlines in ssl_store_load_ca_from_buf() error messages.	2022-11-24 18:45:28 +01:00
Amaury Denoyelle	9875f024ba	BUG/MEDIUM: quic: fix datagram dropping on queueing failed After reading a datagram, it is enqueud for the thread attached to the DCID. This is done via quic_lstnr_dgram_dispatch() function. If this step fails, we remove the datagram from the buffer of quic_receiver_buf. This step is faulty because we use b_del() instead of b_sub(). If quic_receiver_buf was not empty, we will remove content from another datagram while leaving the content of the last read datagram. This probably produces valid datagram dropping and may even result in crash. As stated, this bug can only happen if qc_lstnr_dgram_dispatch() fails which happen on two occaions : * on quic_dgram allocation failure, which should be pretty rare * on datagram labelled as invalid for QUIC protocol. This may happen more frequently depending on the network conditions. Thus, this bug has been labelled as a medium one. This should be backported up to 2.6.	2022-11-24 16:45:02 +01:00
Willy Tarreau	6c72fa3d18	CLEANUP: qpack: properly use the QPACK macros not HPACK ones in debug code There were a few leftovers of DEBUG_HPACK and HPACK_SHT_SIZE instead of their QPACK equivalent in the QPACK debug code. There's no harm anyway, but it could lead to confusing results if the tables are not sized equally.	2022-11-24 15:38:26 +01:00
Willy Tarreau	f87bb23acb	CLEANUP: qpack: fix format string in debugging code (int signedness) In issue #1939, Ilya mentions that cppchecks warned about use of "%d" to report the QPACK table's index that's locally stored as an unsigned int. While technically valid, this will never cause any trouble since indexes are always small positive values, but better use %u anyway to silence this warning.	2022-11-24 15:35:17 +01:00
Willy Tarreau	d05aa38950	CLEANUP: peers: fix format string for status messages (int signedness) In issue #1939, Ilya mentions that cppchecks warned about use of "%d" to report the status state that's locally stored as an unsigned int. While technically valid, this will never cause any trouble since in the end what we store there are the applet's states (just a few enum values). Better use %u anyway to silence this warning.	2022-11-24 15:32:20 +01:00
Aurelien DARRAGON	a7dc251e07	MINOR: auth: silence null dereference warning in check_user() In GH issue #1940 cppcheck warns about a possible null-dereference in check_user() when DEBUG_AUTH is enabled. Indeed, <ep> may potentially be NULL because upon error crypt_r() and crypt() may return a null pointer. However it's not directly derefenced, it is only passed to printf() with '%s' fmt. While it is in practice fine with the printf implementations we care about (that check strings against null before printing them), it is undefined behavior according to the spec, hence the warning. Let's check <ep> before passing it to printf. This should partly solve GH #1940.	2022-11-24 15:24:02 +01:00
Willy Tarreau	95f40c698d	MINOR: sample: make the rand() sample fetch function use the statistical_prng Emeric noticed that producing many randoms to fill a stick table was saturating on the rand_lock. Since 2.4 we have the statistical PRNG for low-quality randoms like this one, there is no point in using the one that was originally implemented for the purpose of creating safe UUIDs, since the doc itself clearly states that these randoms are not secure and they have not been in the past either. With this change, locking contention is completely gone.	2022-11-24 15:04:13 +01:00
Uriah Pollock	3cbf09ed64	MEDIUM: ssl: add minimal WolfSSL support with OpenSSL compatibility mode This adds a USE_OPENSSL_WOLFSSL option, wolfSSL must be used with the OpenSSL compatibility layer. This must be used with USE_OPENSSL=1. WolfSSL build options: ./configure --prefix=/opt/wolfssl --enable-haproxy HAProxy build options: USE_OPENSSL=1 USE_OPENSSL_WOLFSSL=1 WOLFSSL_INC=/opt/wolfssl/include/ WOLFSSL_LIB=/opt/wolfssl/lib/ ADDLIB='-Wl,-rpath=/opt/wolfssl/lib' Using at least the commit 54466b6 ("Merge pull request #5810 from Uriah-wolfSSL/haproxy-integration") from WolfSSL. (2022-11-23). This is still to be improved, reg-tests are not supported yet, and more tests are to be done. Signed-off-by: William Lallemand <wlallemand@haproxy.org>	2022-11-24 11:29:03 +01:00
Willy Tarreau	33a6870fea	BUILD: quic: silence two invalid build warnings at -O1 with gcc-6.5 Gcc 6.5 is now well known for triggering plenty of false "may be used uninitialized", particularly at -O1, and two of them happen in quic, quic_tp and quic_conn. Both of them were reviewed and easily confirmed as wrong (gcc seems to ignore the control flow after the function returns and believes error conditions are not met). Let's just preset the variables that bothers it. In quic_tp the initialization was moved out of the loop since there's no point inflating the code just to silence a stupid warning.	2022-11-24 09:16:41 +01:00
Willy Tarreau	08093cc0fa	CLEANUP: tools: do not needlessly include xxhash nor cli from tools.h These includes brought by commit `9c76637ff` ("MINOR: anon: add new macros and functions to anonymize contents") resulted in an increase of exactly 20% of the number of lines to build. These include are not needed there, only tools.c needs xxhash.h.	2022-11-24 08:30:48 +01:00
Willy Tarreau	07a3d539f5	BUILD: quic: global.h is needed in cfgparse-quic cfgparse-quic accesses some members of the "global" struct but only includes global-t.h. It actually used to work via tools.h due to a long dependency chain that brought it, but it will be fixed and will break cfgparse-quic, so let's fix it first.	2022-11-24 08:30:48 +01:00
Willy Tarreau	a4728584ff	BUILD: stick-tables: fix build breakage in xxhash on older compilers Commit `36d156564` ("MINOR: peers: Support for peer shards") reintroduced a direct dependency on import/xxhash.h which was previously dropped by commit `d5fc8fcb8` ("CLEANUP: Add haproxy/xxhash.h to avoid modifying import/xxhash.h"). This results in xxhash.h being included twice, which breaks the build on older compilers which do not like seeing XXH32_hash_t being defined twice. Let's just use include/haproxy/xxhash.h instead. No backport is needed.	2022-11-24 07:38:13 +01:00
Christopher Faulet	d1b573059a	MINOR: mux-h1: Avoid useless call to h1_send() if no error is sent If we choose to not send an error to the client, there is no reason to call h1_send() for nothing. This happens because functions handling errors return 1 instead of 0 when nothing is performed.	2022-11-23 17:13:13 +01:00
Christopher Faulet	a1a76ce709	MINOR: mux-h1: Remove H1C_F_WAIT_NEXT_REQ in functions handling errors If is cleaner to remove this flag in the internal functions handling errors, responsible to update the H1 connection state, instead to do so in calling functions. This will hopefully avoid bugs in future.	2022-11-23 17:07:49 +01:00
Christopher Faulet	227424450c	BUG/MINOR: mux-h1: Fix handling of 408-Request-Time-Out When a timeout is detected waiting for the request, a 408-Request-Time-Out response is sent. However, an error was introduced by commit 6858d76cd3 ("BUG/MINOR: mux-h1: Obey dontlognull option for empty requests"). Instead of inhibiting the log message, the option was stopping the error sending. Of course, we must do the opposite. This patch must be backported as far as 2.4.	2022-11-23 16:58:23 +01:00
Christopher Faulet	4427ea7f04	BUG/MEDIUM: mux-h1: Remove H1C_F_WAIT_NEXT_REQ flag on a next request When an idle H1 connection starts to process an new request, we must take care to remove H1C_F_WAIT_NEXT_REQ flag. This flag is used to know an idle H1 connection has already processed at least one request and is waiting for a next one, but nothing was received yet. Keeping this flag leads to a crash because some running H1 connections may be erroneously released on a soft-stop. Indeed, only idle or closed connections must be released. This bug was reported into #1903. It is 2.7-specific. No backport needed.	2022-11-23 15:59:00 +01:00
Christopher Faulet	881cce9f13	BUILD: ssl-sock: Silent error about NULL deref in ssl_sock_bind_verifycbk() In ssl_sock_bind_verifycbk(), when compiled without QUIC support, the compiler may report an error during compilation about a possible NULL dereference: src/ssl_sock.c: In function ‘ssl_sock_bind_verifycbk’: src/ssl_sock.c:1738:12: error: potential null pointer dereference [-Werror=null-dereference] 1738 \| ctx->xprt_st \|= SSL_SOCK_ST_FL_VERIFY_DONE; \| ~~~^~~~~~~~~ A BUG_ON() was addeded because it must never happen. But when compiled without DEBUG_STRICT, there is nothing to help the compiler. Thus ALREADY_CHECKED() macro is used. The ssl-sock context and the bind config are concerned. This patch must be backported to 2.6.	2022-11-23 09:27:14 +01:00
Christopher Faulet	92c2de1a06	BUILD: http-htx: Silent build error about a possible NULL start-line In http_replace_req_uri(), if the URI was successfully replaced, it means the start-line exists. However, the compiler reports an error about a possible NULL pointer dereference: src/http_htx.c: In function ‘http_replace_req_uri’: src/http_htx.c:392:19: error: potential null pointer dereference [-Werror=null-dereference] 392 \| sl->flags &= ~HTX_SL_F_NORMALIZED_URI; So, ALREADY_CHECKED() macro is used to silent the build error. This patch must be backported with `84cdbe478a` ("BUG/MINOR: http-htx: Don't consider an URI as normalized after a set-uri action").	2022-11-22 18:02:02 +01:00
Christopher Faulet	a462ee0af4	BUG/MEDIUM: mux-h1: Subscribe for reads on error on sending path The recent refactoring about errors handling in the H1 multiplexer introduced a bug on abort when the client wait for the server response. The bug only exists if abortonclose option is not enabled. Indeed, in this case, when the end of the message is reached, the mux stops to receive data because these data are part of the next request. However, error on the sending path are no longer fatal. An error on the reading path must be caught to do so. So, in case of a client abort, the error is not reported to the upper layer and the H1 connection cannot be closed if some pending data are blocked (remember, an error on sending path was detected, blocking outgoing data). To be sure to have a chance to detect the abort in the case, when an error is detected on the sending path, we force the subscription for reads. This patch, with the previous one, should fix the issue #1943. It is 2.7-specific, no backport is needed.	2022-11-22 17:49:10 +01:00
Christopher Faulet	f75cc5468a	BUG/MEDIUM: mux-h1: Don't release H1C on timeout if there is a SC attached When the H1 task timed out, we must be careful to not release the H1 conneciton if there is still a H1 stream with a stream-connector attached. In this case, we must wait. There are some tests to prevent it happens. But the last one only tests the UPGRADING state while there is also the CLOSING state with a SC attached. But, in the end, it is safer to test if there is a H1 stream with a SC attached. This patch should partially fix the issue #1943. However, it only prevent the segfault. There is another bug under the hood. BTW, this one is 2.7-specific. Not backport is needed.	2022-11-22 17:49:10 +01:00
Christopher Faulet	84cdbe478a	BUG/MINOR: http-htx: Don't consider an URI as normalized after a set-uri action An abosulte URI is marked as normalized if it comes from an H2 client. This way, we know we can send a relative URI to an H1 server. But, after a set-uri action, the URI must no longer be considered as normalized. Otherwise there is no way to send an absolute URI on the server side. If it is important to update a normalized absolute URI without altering this property, the host, path and/or query-string must be set separatly. This patch should fix the issue #1938. It should be backported as far as 2.4.	2022-11-22 17:49:10 +01:00
Christopher Faulet	e16ffb0383	BUG/MINOR: h1: Replace authority validation to conform RFC3986 Except for CONNECT method, where a normalization is performed, we expected to have an exact match between the authority and the host header value. However it was too strict. Indeed, default port must be handled and the matching must respect the RFC3986. There is already a scheme based normalizeation performed on the URI later, on the HTX message. And we cannot normalize the URI during H1 parsing to be able to report proper errors on the original raw buffer. And a systematic read-only normalization to validate the authority will consume CPU for only few requests. At the end, we decided to perform extra-checks when the exact match fails. Now, following authority/host are considered as equivalent: http: domain.com <=> domain.com:80 <=> domain.com: https: domain.com <=> domain.com:443 <=> domain.com: This patch depends on: * MINOR: h1: Consider empty port as invalid in authority for CONNECT * MINOR: http: Considere empty ports as valid default ports It is a bug regarding the RFC3986. Technically, I may be backported as far as 2.4. However, this must be discussed first. If it is backported, the commits above must be backported too.	2022-11-22 17:49:10 +01:00
Christopher Faulet	e5dfe1169d	BUG/MINOR: http-htx: Normalized absolute URIs with an empty port Thanks to the previous commit ("MINOR: http: Considere empty ports as valid default ports"), empty ports are now considered as valid default ports. Thus, absolute URIs with empty port should be normalized. So now, the following URIs are normalized: http://example.com:/ --> http://example.com/ https://example.com:/ --> https://example.com/ This patch depend on: * MINOR: h1: Consider empty port as invalid in authority for CONNECT * MINOR: http: Considere empty ports as valid default ports It is a bug regarding the RFC3986. Technically, I may be backported as far as 2.4. However, this must be discussed first. If backported, the commits above must be backported too.	2022-11-22 16:56:49 +01:00
Christopher Faulet	99ade9e0da	MINOR: http: Considere empty ports as valid default ports In RFC3986#6.2.3, following URIs are considered as equivalent: http://example.com http://example.com/ http://example.com:/ http://example.com:80/ The third one is interristing because the port is empty and it is still considered as a default port. Thus, http_get_host_port() does no longer return IST_NULL when the port is empty. Now, a ist is returned, it points on the first character after the colon (':') with a length of 0. In addition, http_is_default_port() now considers an empty port as a default port, regardless the scheme. This patch must not be backported, if so, without the previous one ("MINOR: h1: Consider empty port as invalid in authority for CONNECT").	2022-11-22 16:56:49 +01:00
Christopher Faulet	75348c2e8b	MINOR: h1: Consider empty port as invalid in authority for CONNECT For now, this change is useless because http_get_host_port() returns IST_NULL when the port is empty. But this will change. For other methods, empty ports are valid. But not for CONNECT method. To still return a 400-Bad-Request if a CONNECT is performed with an empty port, istlen() is used to test the port, instead of isttest().	2022-11-22 16:27:52 +01:00
Aurelien DARRAGON	e3177af465	CLEANUP: tools: extra check in utoa_pad Removing useless check in utoa_pad(). This was reported by Ilya with the help of cppcheck.	2022-11-22 16:27:52 +01:00
Aurelien DARRAGON	ac1ca5cc7b	CLEANUP: arg: remove extra check in make_arg_list arg escaping Len cannot be equal to 1 when entering in escape handling code. But yet, an extra "len == 1" check was performed. Removing this useless check. This was reported by Ilya with the help of cppcheck.	2022-11-22 16:27:52 +01:00
Aurelien DARRAGON	ab9efc25f0	BUG/MINOR: log: fix parse_log_message rfc5424 size check In parse_log_message(), if log is rfc5424 compliant, p pointer is incremented and size is not. However size is still used in further checks as if p pointer was not incremented. This could lead to logic error or buffer overflow if input buf is not null-terminated. Fixing this by making sure size is up to date where it is needed. It could be backported up to 2.4.	2022-11-22 16:27:52 +01:00
Aurelien DARRAGON	9dce88ba2c	BUG/MINOR: cfgparse-listen: fix ebpt_next_dup pointer dereference on proxy "from" inheritance ebpt_next_dup() was used 2 times in a row but only the first call was checked against NULL, probably assuming that the 2 calls always yield the same result here. gcc is not OK with that, and it should be safer to store the result of the first call in a temporary var to dereference it once checked against NULL. This should fix GH #1869. Thanks to Ilya for reporting this issue. It may be backported up to 2.4.	2022-11-22 16:27:52 +01:00
Willy Tarreau	5ec79f1a04	BUILD: sched: fix build with DEBUG_THREAD with the previous commit The build with DEBUG_THREAD was broken by commit `fc50b9dd1` ("BUG/MAJOR: sched: protect task during removal from wait queue"). It took me a while to figure how to declare and aligned and initialized rwlock that wasn't static, but it turns out that __decl_aligned_rwlock() does exactly this, so that we don't have to assign an integer value when a struct is expected in case of debugging. No backport is needed.	2022-11-22 10:24:07 +01:00
Willy Tarreau	fc50b9dd14	BUG/MAJOR: sched: protect task during removal from wait queue The issue addressed by commit `fbb934da9` ("BUG/MEDIUM: stick-table: fix a race condition when updating the expiration task") is still present when thread groups are enabled, but this time it lies in the scheduler. What happens is that a task configured to run anywhere might already have been queued into one group's wait queue. When updating a stick table entry, sometimes the task will have to be dequeued and requeued. For this a lock is taken on the current thread group's wait queue lock, but while this is necessary for the queuing, it's not sufficient for dequeuing since another thread might be in the process of expiring this task under its own group's lock which is different. This is easy to test using 3 stick tables with 1ms expiration, 3 track-sc rules and 4 thread groups. The process crashes almost instantly under heavy traffic. One approach could consist in storing the group number the task was queued under in its descriptor (we don't need 32 bits to store the thread id, it's possible to use one short for the tid and another one for the tgrp). Sadly, no safe way to do this was figured, because the race remains at the moment the thread group number is checked, as it might be in the process of being changed by another thread. It seems that a working approach could consist in always having it associated with one group, and only allowing to change it under this group's lock, so that any code trying to change it would have to iterately read it and lock its group until the value matches, confirming it really holds the correct lock. But this seems a bit complicated, particularly with wait_expired_tasks() which already uses upgradable locks to switch from read state to a write state. Given that the shared tasks are not that common (stick-table expirations, rate-limited listeners, maybe resolvers), it doesn't seem worth the extra complexity for now. This patch takes a simpler and safer approach consisting in switching back to a single wq_lock, but still keeping separate wait queues. Given that shared wait queues are almost always empty and that otherwise they're scanned under a read lock, the contention remains manageable and most of the time the lock doesn't even need to be taken since such tasks are not present in a group's queue. In essence, this patch reverts half of the aforementionned patch. This was tested and confirmed to work fine, without observing any performance degradation under any workload. The performance with 8 groups on an EPYC 74F3 and 3 tables remains twice the one of a single group, with the contention remaining on the table's lock first. No backport is needed.	2022-11-22 09:10:08 +01:00
Willy Tarreau	469fa47950	BUILD: listener: fix build warning on global_listener_rwlock without threads The global_listener_rwlock was introduced by recent commit `13e86d947` ("BUG/MEDIUM: listener: Fix race condition when updating the global mngmt task"), but it's declared static and is not used when threads are disabled, thus causing a warning to be emitted in this case. Let's just condition it to thread usage to shut the warning. This will need to be backported where the patch above is backported.	2022-11-22 09:10:08 +01:00
Willy Tarreau	c21a187ec0	MINOR: server/idle: make the next_takeover index per-tgroup In order to evenly pick idle connections from other threads, there is a "next_takeover" index in the server, that is incremented each time a connection is picked from another thread, and indicates which one to start from next time. With thread groups this doesn't work well because the index is the same regardless of the group, and if a group has more threads than another, there's even a risk to reintroduce an imbalance. This patch introduces a new per-tgroup storage in servers which, for now, only contains an instance of this next_takeover index. This way each thread will now only manipulate the index specific to its own group, and the takeover will become fair again. More entries may come soon.	2022-11-21 19:21:07 +01:00
Willy Tarreau	9dc231a6b2	BUG/MINOR: server/idle: at least use atomic stores when updating max_used_conns In 2.2, some idle conns usage metrics were added by commit `cf612a045` ("MINOR: servers: Add a counter for the number of currently used connections."), which mentioned that the operation doesn't need to be atomic since we're not seeking exact values. This is true but at least we should use atomic stores to make sure not to cause invalid values to appear on archs that wouldn't guarantee atomicity when writing an int, such as writing two 16-bit words. This is pretty unlikely on our targets but better keep the code safe against this. This may be backported as far as 2.2.	2022-11-21 19:21:07 +01:00
Willy Tarreau	fdecaf6ae4	BUG/MINOR: resolvers: do not run the timeout task when there's no resolution The function resolv_update_resolvers_timeout() always schedules a wakeup of the process_resolvers() task based on the "timeout resolve" setting, regardless of the presence of an ongoing resolution or not. This is causing one wakeup every second by default even when there's no resolvers section (due to the default one), and can even be worse: creating a section with "timeout resolve 1" bombs the process with 1000 wakeups per second. Let's condition the setting to the presence of a resolution to address this. This issue has been there forever, but it doesn't cause that much trouble, and given how fragile and tricky this code is, it's probably wise to refrain from backporting it until it's reported to really cause trouble.	2022-11-21 19:21:07 +01:00
Amaury Denoyelle	28ea31c7cb	MINOR: global: generate random cluster.secret if not defined If no cluster-secret is defined by the user, a random one is silently generated. This ensures that at least QUIC Retry tokens are generated if abnormal conditions are detected. However, it is advisable to specify it in the configuration for tokens to be valid even after a reload or across LBs instances in the same cluster. This should be backported up to 2.6.	2022-11-21 16:41:34 +01:00
Amaury Denoyelle	996ca7d0fa	MINOR: quic: report error if force-retry without cluster-secret QUIC Retry generation relies on global cluster-secret to produce token valid even after a process restart and across several LBs instances. Before this patch, Retry is automatically deactivated if no cluster-secret is provided. This is the case even if a user has configured a QUIC listener with quic-force-retry. Change this behavior by now returning an error during configuration parsing. The user must provide a cluster-secret if quic-force-retry is used. This shoud be backported up to 2.6.	2022-11-21 16:34:09 +01:00
Willy Tarreau	7583c36790	MINOR: cli/pools: add pool name filtering capability to "show pools" Now it becomes possible to match a pool name's prefix, for example: $ socat - /tmp/haproxy.sock <<< "show pools match quic byusage" Dumping pools usage. Use SIGQUIT to flush them. - Pool quic_conn_r (65560 bytes) : 1337 allocated (87653720 bytes), ... - Pool quic_crypto (1048 bytes) : 6685 allocated (7005880 bytes), ... - Pool quic_conn (4056 bytes) : 1337 allocated (5422872 bytes), ... - Pool quic_rxbuf (262168 bytes) : 8 allocated (2097344 bytes), ... - Pool quic_connne (184 bytes) : 9359 allocated (1722056 bytes), ... - Pool quic_frame (184 bytes) : 7938 allocated (1460592 bytes), ... - Pool quic_tx_pac (152 bytes) : 6454 allocated (981008 bytes), ... - Pool quic_tls_ke (56 bytes) : 12033 allocated (673848 bytes), ... - Pool quic_rx_pac (408 bytes) : 1596 allocated (651168 bytes), ... - Pool quic_tls_se (88 bytes) : 6685 allocated (588280 bytes), ... - Pool quic_cstrea (88 bytes) : 4011 allocated (352968 bytes), ... - Pool quic_tls_iv (24 bytes) : 12033 allocated (288792 bytes), ... - Pool quic_dgram (344 bytes) : 732 allocated (251808 bytes), ... - Pool quic_arng (56 bytes) : 4011 allocated (224616 bytes), ... - Pool quic_conn_c (152 bytes) : 1337 allocated (203224 bytes), ... Total: 15 pools, 109578176 bytes allocated, 109578176 used ... In this case the reported total only concerns the dumped ones.	2022-11-21 10:14:52 +01:00
Willy Tarreau	2fba08faec	MINOR: cli/pools: add sorting capabilities to "show pools" The "show pools" command is used a lot for debugging but didn't get much love over the years. This patch brings new capabilities: - sorting the output by pool names to ese their finding ("byname"). - sorting the output by reverse item size to spot the biggest ones("bysize") - sorting the output by reverse number of allocated bytes ("byusage") The last one (byusage) also omits displaying the ones with zero allocation. In addition, an optional max number of output entries may be passed so as to dump only the N most relevant ones.	2022-11-21 10:14:52 +01:00
Willy Tarreau	224adf2bfb	MINOR: cli/pools: store "show pools" results into a temporary array This will permit sorting and filtering that are currently not possible.	2022-11-21 10:14:52 +01:00
Ilya Shipitsin	ace3da8dd4	CLEANUP: quic: replace "choosen" with "chosen" all over the code Some variables were set as "choosen" instead of "chosen", this is dedicated spelling fix	2022-11-21 09:22:28 +01:00
Frédéric Lécaille	74b5f7b31b	BUG/MAJOR: quic: Crash after discarding packet number spaces This previous patch was not sufficient to prevent haproxy from crashing when some Handshake packets had to be inspected before being possibly retransmitted: "BUG/MAJOR: quic: Crash upon retransmission of dgrams with several packets" This patch introduced another issue: access to packets which have been released because still attached to others (in the same datagram). This was the case for instance when discarding the Initial packet number space before inspecting an Handshake packet in the same datagram through its ->prev or member in our case. This patch implements quic_tx_packet_dgram_detach() which detaches a packet from the adjacent ones in the same datagram to be called when ackwowledging a packet (as done in the previous commit) and when releasing its memory. This was, we are sure the released packets will not be accessed during retransmissions. Thank you to @gabrieltz for having reported this issue in GH #1903. Must be backported to 2.6.	2022-11-20 18:35:46 +01:00
Abhijeet Rastogi	c8601507b2	MINOR: cli: print parsed command when not found It is useful because when we're passing data to runtime API, specially via code, we can mistakenly send newlines leading to some lines being wrongly interpretted as commands. This is analogous to how it's done in a shell, example bash: $ not_found arg1 bash: not_found: command not found... $ Real world example: Following the official docs to add a cert: $ echo -e "set ssl cert ./cert.pem <<\n$(cat ./cert.pem)\n" \| socat stdio tcp4-connect:127.0.0.1:9999 Note, how the payload is sent via '<<\n$(cat ./cert.pem)\n'. If cert.pem contains a newline between various PEM blocks, which is valid, the above command would generate a flood of 'Unknown command' messages for every line sent after the first newline. As a new user, this detail is not clearly visible as socket API doesn't say what exactly what was 'unknown' about it. The cli interface should be obvious around guiding user on "what do do next". This commit changes that by printing the parsed cmd in output like 'Unknown command: "<cmd>"' so the user gets clear "next steps", like bash, regarding what indeed was the wrong command that HAproxy couldn't interpret. Previously: $ echo -e "show version\nhelpp"\| socat ./haproxy.sock - \| head -n4 2.7-dev6 Unknown command, but maybe one of the following ones is a better match: add map [@<ver>] <map> <key> <val> : add a map entry (payload supported instead of key/val) Now: $ echo -e "show version\nhelpp"\| socat ./haproxy.sock - \| head -n4 2.7-dev8-737bb7-156 Unknown command: 'helpp', but maybe one of the following ones is a better match: add map [@<ver>] <map> <key> <val> : add a map entry (payload supported instead of key/val)	2022-11-19 05:04:49 +01:00
Fr�d�ric L�caille	814645f42f	BUG/MAJOR: quic: Crash upon retransmission of dgrams with several packets As revealed by some traces provided by @gabrieltz in GH #1903 issue, there are clients (chrome I guess) which acknowledge only one packet among others in the same datagram. This is the case for the first datagram sent by a QUIC haproxy listener made an Initial packet followed by an Handshake one. In this identified case, this is the Handshake packet only which is acknowledged. But if the client is able to respond with an Handshake packet (ACK frame) this is because it has successfully parsed the Initial packet. So, why not also acknowledging it? AFAIK, this is mandatory. On our side, when restransmitting this datagram, the Handshake packet was accessed from the Initial packet after having being released. Anyway. There is an issue on our side. Obviously, we must not expect an implementation to respect the RFC especially when it want to build an attack ;) With this simple patch for each TX packet we send, we also set the previous one in addition to the next one. When a packet is acknowledged, we detach the next one and the next one in the same datagram from this packet, so that it cannot be resent when resending these packets (the previous one, in our case). Thank you to @gabrieltz for having reported this issue. Must be backported to 2.6.	2022-11-19 04:56:55 +01:00
Mathias Weiersmueller	d9b7174d99	MEDIUM: tcp-act: add parameter rst-ttl to silent-drop The silent-drop action was extended with an additional optional parameter, [rst-ttl <ttl> ], causing HAProxy to send a TCP RST with the specified TTL towards the client. With this behaviour, the connection state on your own client- facing middle-boxes (load balancers, firewalls) will be purged, but the client will still assume the TCP connection is up because the TCP RST packet expires before reaching the client.	2022-11-19 04:53:47 +01:00
Amaury Denoyelle	2f668f0e60	MINOR: quic: complete traces/debug for handshake Add more traces to follow CRYPTO data buffering in ncbuf. Offset for quic_enc_level is now reported for event QUIC_EV_CONN_PRHPKTS. Also ncb_advance() must never fail so a BUG_ON() statement is here to guarantee it. This was useful to track handshake failure reported too often. This is related to github issue #1903. This should be backported up to 2.6.	2022-11-18 16:44:46 +01:00
Amaury Denoyelle	bc174b2101	BUG/MEDIUM: quic: fix memleak for out-of-order crypto data Liberate quic_enc_level ncbuf in quic_stream_free(). In most cases, this will already be done when handshake is completed via qc_treat_rx_crypto_frms(). However, if a connection is released before handshake completion, a leak was present without this patch. Under normal situation, this leak should have been limited due to the majority of QUIC connection success on handshake. However, another bug caused handshakes to fail too frequently, especially with chrome client. This had the side-effect to dramatically increase this memory leak. This should fix in part github issue #1903.	2022-11-18 16:44:46 +01:00
Amaury Denoyelle	ff95f2d447	BUG/MEDIUM: quic: fix unsuccessful handshakes on ncb_advance error QUIC handshakes were frequently in error due to haproxy misuse of ncbuf. This resulted in one of the following scenario : - handshake rejected with CONNECTION_CLOSE due to overlapping data rejected - CRYPTO data fully received by haproxy but handshake completion signal not reported causing the client to emit PING repeatedly before timeout This was produced because ncb_advance() result was not checked after providing crypto data to the SSL stack in qc_provide_cdata(). However, this operation can fail if a too small gap is formed. In the meantime, quic_enc_level offset was always incremented. In some cases, this caused next ncb_add() to report rejected overlapping data. In other cases, no error was reported but SSL stack never received the end of CRYPTO data. Change slightly the handling of new CRYPTO frames to avoid this bug : when receiving a CRYPTO frame for the current offset, handle them directly as previously done only if quic_enc_level ncbuf is empty. In the other case, copy them to the buffer before treating contiguous data via qc_treat_rx_crypto_frms(). This change ensures that ncb_advance() operation is now conducted only in a data block : thus this is guaranteed to never fail. This bug was easily reproduced with chromium as it fragments CRYPTO frames randomly in several frames out of order. This commit has two drawbacks : - it is slightly less worst on performance because as sometimes even data at the current offset will be memcpy - if a client uses too many fragmented CRYPTO frames, this can cause repeated ncb_add() error on gap size. This can be reproduced with chrome, albeit with a slighly less frequent rate than the fixed issue. This change should fix in part github issue #1903. This must be backported up to 2.6.	2022-11-18 16:44:46 +01:00
Amaury Denoyelle	7f0295f08a	MINOR: ncbuf: complete doc for ncb_advance() ncb_advance() operation may reject the operation if a too small gap is formed in buffer front. This must be documented to avoid an issue with it. This should be backported up to 2.6.	2022-11-18 16:44:46 +01:00
Christopher Faulet	4cfdcbbd19	BUILD: peers: Remove unused variables Since `0909f62266` ("BUG/MEDIUM: peers: messages about unkown tables not correctly ignored"), the 'sc' variable is no longer used in peer_treat_updatemsg() and peer_treat_definemsg() functions. So, we must remove them to avoid compilation warning. This patch must be backported with the commit above.	2022-11-18 16:40:56 +01:00
Christopher Faulet	5534334f1f	MEDIUM: thread: Restric nbthread/thread-group(s) to very first global sections nbhread, thead-group and thread-groups directives must only be defined in very first global sections. It means no other section must have been parsed before. Indeed, some parts of the configuratio depends on the value of these settings and it is undefined to change them after.	2022-11-18 16:03:45 +01:00
Christopher Faulet	037e3f8735	MINOR: cfgparse: Always check the section position In diag mode, the section position is checked and a warning is emitted if a global section is defined after any non-global one. Now, this check is always performed. But the warning is still only emitted in diag mode. In addition, the result of this check is now stored in a global variable, to be used from anywhere. The aim of this patch is to be able to restrict usage of some global directives to the very first global sections. It will be useful to avoid undefined behaviors. Indeed, some config parts may depend on global settings and it is a problem if these settings are changed after.	2022-11-18 16:03:45 +01:00
Emeric Brun	0909f62266	BUG/MEDIUM: peers: messages about unkown tables not correctly ignored Table defintion's messages and update messages are not correctly ignored if the table is not configured on the local peer. It is a bug because, receiving those messages, the parser returns an error and the upper layer considers that the state of the peer's connection is modified (as it is done in the case of protocol error) and switch immediatly the automate to process the new state. But, even if message is silently ignored because the connection's state doesn't change and we continue to process the next message, some processing remains not performed: for instance the ALIVE flag is not set on the peer's connection as it should be done after receiving any valid messages. This results in a shutdown of the connection when timeout is elapsed as if no message has been received during this delay. This patch fix the behavior, those messages are now silently ignored and the upper layer continue the processing as it is done for any valid messages. This bug appears with the code re-work of the peers on 2.0 so it should be backported until this version.	2022-11-18 15:54:33 +01:00
William Lallemand	b60a77b6d0	BUG/MINOR: ssl: don't initialize the keylog callback when not required The registering of the keylog callback seems to provoke a loss of performance. Disable the registration as well as the fetches if tune.ssl.keylog is off. Must be backported as far as 2.2.	2022-11-18 15:24:23 +01:00
Christopher Faulet	dfefebcd7a	BUG/MEDIUM: raw-sock: Don't report connection error if something was received In raw_sock_to_buf(), if a low-level error is reported, we no longer immediately set an error on the connexion if something was received. This may happen when a RST is received with data. This way, we let a chance to the mux to process received data first instead of immediately aborting. This patch should fix some spurious health-check failures. It is pretty hard to observe, but with a server immediately returning the response followed by a RST, without waiting the request, it is possible to have some health-check errors. For instance, with the following tcploop server: tcploop 8000 L Q W N1 A S:"HTTP/1.0 200 OK\r\n\r\n" F K ( Accept -> send response -> FIN -> Close) we can have such strace output: 15:11:21.433005 socket(AF_INET, SOCK_STREAM, IPPROTO_IP) = 38 15:11:21.433141 fcntl(38, F_SETFL, O_RDONLY\|O_NONBLOCK) = 0 15:11:21.433233 setsockopt(38, SOL_TCP, TCP_NODELAY, [1], 4) = 0 15:11:21.433359 setsockopt(38, SOL_TCP, TCP_QUICKACK, [0], 4) = 0 15:11:21.433457 connect(38, {sa_family=AF_INET, sin_port=htons(8000), sin_addr=inet_addr("127.0.0.1")}, 16) = -1 EINPROGRESS (Operation now in progress) 15:11:21.434215 epoll_ctl(4, EPOLL_CTL_ADD, 38, {events=EPOLLIN\|EPOLLOUT\|EPOLLRDHUP, data={u32=38, u64=38}}) = 0 15:11:21.434468 epoll_wait(4, [{events=EPOLLOUT, data={u32=38, u64=38}}], 200, 21) = 1 15:11:21.434810 recvfrom(38, 0x7f32a83e5020, 16320, 0, NULL, NULL) = -1 EAGAIN (Resource temporarily unavailable) 15:11:21.435405 sendto(38, "OPTIONS / HTTP/1.0\r\ncontent-leng"..., 41, MSG_DONTWAIT\|MSG_NOSIGNAL, NULL, 0) = 41 15:11:21.435833 epoll_ctl(4, EPOLL_CTL_MOD, 38, {events=EPOLLIN\|EPOLLRDHUP, data={u32=38, u64=38}}) = 0 15:11:21.435907 epoll_wait(4, [{events=EPOLLIN\|EPOLLERR\|EPOLLHUP\|EPOLLRDHUP, data={u32=38, u64=38}}], 200, 17) = 1 15:11:21.436024 recvfrom(38, "HTTP/1.0 200 OK\r\n\r\n", 16320, 0, NULL, NULL) = 19 15:11:21.436189 close(38) = 0 15:11:21.436402 write(2, "[WARNING] (163564) : Server bac"..., 184[WARNING] (163564) : Server back-http/www is DOWN, reason: Socket error, check duration: 5ms. 0 active and 0 backup servers left. 0 sessions active, 0 requeued, 0 remaining in queue. The response was received, but it is ignored because an error was reported too. The error handling must be refactored. But it a titanic stain. Thus, for now, a good fix is to delay the error report when something was received. The error will be reported on the next receive, if any. This patch should fix the issue #1863, but it must be confirmed. At least it fixes the above example. It must be backported to 2.6. For older versions, it must be evaluated first.	2022-11-18 15:12:23 +01:00
Aurelien DARRAGON	5ad2b64262	BUG/MINOR: http_ana/txn: don't re-initialize txn and req var lists In http_create_txn(): vars_init_head() was performed on both s->vars_txn and s->var_reqres lists. But this is wrong, these two lists are already initialized upon stream creation in stream_new(). Moreover, between stream_new() and http_create_txn(), some variable may be defined (e.g.: by the frontend), resulting in lists not being empty. Because of this "extra" list initialization, already defined variables can be lost. This causes txn dependant code not being able to access previously defined variables as well as memory leak because http_destroy_txn() relies on these lists to perform the purge. This proved to be the case when a frontend sets variables and lua sample fetch is used in backend section as described in GH #1935. Many thanks to Darragh O'Toole for his detailed report. Removing extra var_init_head (x2) in http_create_txn() to fix the issue. Adding somme comments in the code in an attempt to prevent future misuses of s->var_reqres, and s->var_txn lists. It should be backported in every stable version. (This is an old bug that seems to exist since 1.6-dev6) [cf: On 2.0 and 1.8, for the legacy HTTP code, vars_init() are used during the TXN cleanup, when the stream is reused. So, these calls must be moved from http_init_txn() to http_reset_txn() and not removed.]	2022-11-18 10:20:44 +01:00
Christopher Faulet	ce7928d19c	CLEANUP: mux-h1: Don't test h1c in h1_shutw_conn() The H1 connection cannot be NULL when h1_shutw_conn() is called. Thus there is no reason to test it. This patch should fix the issue #1936.	2022-11-18 08:44:46 +01:00
Christopher Faulet	e6ef4cd747	BUG/MINOR: mux-h1: Fix error handling when H1S allocation failed on client side The goto label is not at the right place. When H1S allocation failed, the error is immediately handled. Thus, "no_parsing" label must be set just after h1_send() call to skip the request parsing part. It is 2-7-specific. No backport needed.	2022-11-17 15:54:13 +01:00
Christopher Faulet	ddfb50eec6	CLEANUP: listener: Remove useless task_queue from manage_global_listener_queue At the end of manage_global_listener_queue(), the task expire date is set to TICK_ETERNITY. Thus, it is useless to call task_queue() just after because the function does nothing in this case.	2022-11-17 15:18:59 +01:00
Christopher Faulet	13e86d947d	BUG/MEDIUM: listener: Fix race condition when updating the global mngmt task It is pretty similar to `fbb934da90` ("BUG/MEDIUM: stick-table: fix a race condition when updating the expiration task"). When the global management task is running, at the end of its process function, it resets the expire date by setting it to TICK_ETERNITY. In same time, a listener may reach a global limit and decides to schedule the task. Thus it is possible to queue the task and trigger the BUG_ON() on the expire date because its value was set to TICK_ETERNITY in the means time: FATAL: bug condition "task->expire == 0" matched at src/task.c:310 call trace(12): \| 0x662de8 [b8 01 00 00 00 c6 00 00]: __task_queue+0xc7/0x11e \| 0x63b03f [48 b8 04 00 00 00 05 00]: main+0x2535e \| 0x63ed1a [e9 d2 fd ff ff 48 8b 45]: listener_accept+0xf72/0xfda \| 0x6a36d3 [eb 01 90 c9 c3 55 48 89]: sock_accept_iocb+0x82/0x87 \| 0x6af22f [48 8b 05 ca f9 13 00 8b]: fd_update_events+0x35a/0x497 \| 0x42a7a8 [89 45 d8 83 7d d8 02 75]: main-0x1eb539 \| 0x6158fb [48 8b 05 e6 06 1c 00 64]: run_poll_loop+0x2e7/0x319 \| 0x615b6c [48 8b 05 ed 65 1d 00 48]: main-0x175 \| 0x7ffff775bded [e9 69 fe ff ff 48 8b 4c]: libc:+0x8cded \| 0x7ffff77e1370 [48 89 c7 b8 3c 00 00 00]: libc:+0x112370 To fix the bug, a RW lock is introduced. It is used to fix the race condition. A read lock is taken when the task is scheduled, in listener_accpet() and a write lock is used at the end of process function to set the expire date to TICK_ETERNITY. This lock should not be used very often and most of time by "readers". So, the impact should be really limited. This patch should fix the issue #1930. It must be backported as far as 1.8 with some cautions because the code has evolved a lot since then.	2022-11-17 15:18:40 +01:00
Christopher Faulet	62138aab3e	MINOR: mux-h1: Rely on a H1S flag to know a WS key was found or not h1_process_mux() is written to allow partial headers formatting. For now, all headers are forwarded in one time. But it is still good to keep this ability at the H1 mux level. So we must rely on a H1S flag instead of a local variable to know a WebSocket key was found in headers to be able to generate a key if necessary. There is no reason to backport this patch.	2022-11-17 14:33:15 +01:00
Christopher Faulet	7f6aa56d47	MINOR: sconn: Set SE_FL_ERROR only when there is no more data to read SE_FL_ERR_PENDING flag is used when there is still data to be read. So we must take care to not set SE_FL_ERROR too early. Thus, on sending path, it must be set if SE_FL_EOS was already set.	2022-11-17 14:33:15 +01:00
Christopher Faulet	ab79b321d6	MEDIUM: mux-fcgi: Introduce flags to deal with connection read/write errors Similarly to the H1 and H2 multiplexers, FCFI_CF_ERR_PENDING is now used to report an error when we try to send data and FCGI_CF_ERROR to report an error when we try to read data. In other funcions, we rely on these flags instead of connection ones. Only FCGI_CF_ERROR is considered as a final error. FCGI_CF_ERR_PENDING does not block receive attempt. In addition, FCGI_CF_EOS flag was added. we rely on it to test if a read0 was received or not.	2022-11-17 14:33:15 +01:00
Christopher Faulet	68ee7845cf	CLEANUP: mux-h2: Remove unused fields in h2c structures Some fields in h2c structures are not used: .mfl, .mft and .mff. Just remove them. .msi field is also removed. It is tested but never set, except when a H2 connection is initialized. It also means h2c_mux_busy() function is useless because it always returns 0 (.msi is always -1). And thus, by transitivity, H2_CF_DEM_MBUSY is also useless because it is never set. So .msi field, h2c_mux_busy() function and H2C_MUX_BUSY flag are removed.	2022-11-17 14:33:15 +01:00
Christopher Faulet	ff7925dce0	MEDIUM: mux-h2: Introduce flags to deal with connection read/write errors Similarly to the H1 multiplexer, H2_CF_ERR_PENDING is now used to report an error when we try to send data and H2_CF_ERROR to report an error when we try to read data. In other funcions, we rely on these flags instead of connection ones. Only H2_CF_ERROR is considered as a final error. H2_CF_ERR_PENDING does not block receive attempt. In addition, we rely on H2_CF_RCVD_SHUT flag to test if a read0 was received or not.	2022-11-17 14:33:15 +01:00
Christopher Faulet	b65af26e19	MEDIUM: mux-pt: Don't always set a final error on SE on the sending path SE_FL_ERROR must be set on the SE descriptor only if EOS was already reported. So call se_fl_set_error() function to properly the ERR_PENDING/ERROR flags. It is not really a bug because the mux-pt is really simple. But it is better to do it now the right way.	2022-11-17 14:33:15 +01:00
Christopher Faulet	31da34d1e7	MEDIUM: mux-h1: Don't report a final error whe a message is aborted When the H1 connection is aborted, we no longer set a final error. To do so, the flag H1C_F_ABORTED was added. For now, it is only set when a error is detected on the H1 stream. Idea is to use ERR_PENDING/ERROR for upgoing errors and ABRT_PENDING/ABRTED for downgoing errors.	2022-11-17 14:33:15 +01:00
Christopher Faulet	fc473a6453	MEDIUM: mux-h1: Rely on the H1C to deal with shutdown for reads read0 is now handled with a H1 connection flag (H1C_F_EOS). Corresponding flag was removed on the H1 stream and we fully rely on the SE descriptor at the stream level. Concretly, it means we rely on the H1 connection flags instead of the connection one. H1C_F_EOS is only set in h1_recv() or h1_rcv_pipe() after a read if a read0 was detected.	2022-11-17 14:33:15 +01:00
Christopher Faulet	bef8900cd6	MINOR: mux-h1: Add flag on H1 stream to deal with internal errors A new error is added on H1 stream to deal with internal errors. For now, this error is only reported when we fail to create a stream-connector. This way, the error is reported at the H1 stream level and not the H1 connection level.	2022-11-17 14:33:14 +01:00
Christopher Faulet	56a499475f	CLEANUP: mux-h1: Rename H1C_F_ERR_PENDING into H1C_F_ABRT_PENDING H1C_F_ERR_PENDING flags will be used to refactor error handling at the H1 connection level. It will be used to notify error during sends. Thus, the flag to notify an error must be sent before closing the connection is now named H1C_F_ABRT_PENDING. This introduce a naming convertion: ERROR must be used to notify upper layer of an event at the lower ones while ABORT must be used in the opposite direction.	2022-11-17 14:33:14 +01:00
Christopher Faulet	2177d96acd	MINOR: mux-h1: Don't handle subscribe for reads in h1_process_demux() When the request headers are not fully received, we must subscribe the H1 connection for reads to be able to receive more data. This was performed in h1_process_demux(). It is now perfoemd in h1_process_demux().	2022-11-17 14:33:14 +01:00
Christopher Faulet	4e72b172d7	MEDIUM: mux-h1: Handle H1C states via its state field instead of H1C_F_ST_* The H1 connection state is now handled in a dedicated state. H1C_F_ST_* flags are removed. All states are now exclusives. It is easier to know the H1 connection states. It is alive, or usable, if it is not CLOSING or CLOSED. It is CLOSING if it should be closed ASAP but a stream is still attached and/or the output buffer is not empty. CLOSED is used when the H1 connection is ready to be closed. Other states are quite easy to understand. There is no special changes in the H1 connection behavior. Except in h1_send(). When a CLOSING connection is CLOSED, the function now reports an activity. In addition, when an embryonic H1 stream is aborted, it is destroyed. This way, the H1 connection can be switched to CLOSED state.	2022-11-17 14:33:14 +01:00
Christopher Faulet	ef93be2a7b	MINOR: mux-h1: Add a dedicated enum to deal with H1 connection state The H1 connection state will be handled is a dedicated field. To do so, h1_cs enum was added. The different states are more or less equivalent to H1C_F_ST_* flags: * H1_CS_IDLE <=> H1C_F_ST_IDLE * H1_CS_EMBRYONIC <=> H1C_F_ST_EMBRYONIC * H1_CS_UPGRADING <=> H1C_F_ST_ATTACHED && !H1C_F_ST_READY * H1_CS_RUNNING <=> H1C_F_ST_ATTACHED && H1C_F_ST_READY * H1_CS_CLOSING <=> H1C_F_ST_SHUTDOWN && (H1C_F_ST_ATTACHED \|\| b_data(&h1c->ibuf)) * H1_CS_CLOSED <=> H1C_F_ST_SHUTDOWN && !H1C_F_ST_ATTACHED && !b_data(&h1c->ibuf) In addition, in this patch, the h1_is_alive() and h1_close() function are added. The first one will be used to know if a H1 connection is alive or not. The second one will be used to set the connection in CLOSING or CLOSED state, depending on the output buffer state and if there is still a H1 stream or not. For now, the H1 connection state is not used.	2022-11-17 14:33:14 +01:00
Christopher Faulet	71abc0cfd5	CLEANUP: mux-h1: Rename H1C_F_ST_ERROR and H1C_F_ST_SILENT_SHUT flags _ST_ part is removed from these 2 flags because they don't reflect a state. In addition, the H1 connection state will be handled in a dedicated enum.	2022-11-17 14:33:14 +01:00
Christopher Faulet	089cc6e805	REORG: mux-h1: Reorg the H1C structure Fields in H1C structure are reorganised to not have the output buffer straddled between to cache lines. There is 4-bytes hole after the flags, but it will be partially filled by an enum representing the H1 connection state.	2022-11-17 14:33:14 +01:00
Christopher Faulet	7fcbcc0e4c	CLEANUP: mux-h1; Rename H1S_F_ERROR flag into H1S_F_ERROR_MASK In fact, H1S_F_ERROR is not a flag but a mask. So rename it to make it clear.	2022-11-17 14:33:14 +01:00
Christopher Faulet	c3fe6f3b7a	MINOR: mux-h1: Remove usless code inside shutr callback For now, at the transport-layer level, there is no shutr callback (in xprt_ops). Thus, to ease the aborts refacotring, the code is removed from the h1_shutr() function. The callback is not removed for now. It is only kept to have a trace message. It may be handy for debugging sessions.	2022-11-17 14:33:14 +01:00
Willy Tarreau	0c5e9896c7	BUG/MINOR: pool/cli: use ullong to report total pool usage in bytes As noticed by Gabriel Tzagkarakis in issue #1903, the total pool size in bytes is historically still in 32 bits, but at least we should report the product of the number of objects and their size in 64 bits so that the value doesn't wrap around 4G. This may be backported to all versions.	2022-11-17 11:10:53 +01:00
Amaury Denoyelle	3a72ba2aed	BUILD: quic: fix dubious 0-byte overflow on qc_release_lost_pkts With GCC 12.2.0 and O2 optimization activated, compiler reports the following warning for qc_release_lost_pkts(). In function ‘quic_tx_packet_refdec’, inlined from ‘qc_release_lost_pkts.constprop’ at src/quic_conn.c:2056:3: include/haproxy/atomic.h:320:41: error: ‘__atomic_sub_fetch_4’ writing 4 bytes into a region of size 0 overflows the destination [-Werror=stringop-overflow=] 320 \| #define HA_ATOMIC_SUB_FETCH(val, i) __atomic_sub_fetch(val, i, __ATOMIC_SEQ_CST) \| ^~~~~~~~~~~~~~~~~~ include/haproxy/quic_conn.h:499:14: note: in expansion of macro ‘HA_ATOMIC_SUB_FETCH’ 499 \| if (!HA_ATOMIC_SUB_FETCH(&pkt->refcnt, 1)) { \| ^~~~~~~~~~~~~~~~~~~ GCC thinks that quic_tx_packet_refdec() can be called with a NULL argument from qc_release_lost_pkts() with <oldest_lost> as arg. This warning is a false positive as <oldest_lost> cannot be NULL in qc_release_lost_pkts() at this stage. This is due to the previous check to ensure that <pkts> list is not empty. This warning is silenced by using ALREADY_CHECKED() macro. This should be backported up to 2.6. This should fix github issue #1852.	2022-11-17 10:34:47 +01:00
Willy Tarreau	1b662aabbf	BUG/MEDIUM: ring: fix creation of server in uninitialized ring If a "ring" section initialization fails (e.g. due to a duplicate name, invalid chars, or missing memory), any subsequent "server" statement that appears in the same section will crash the config parser by dereferencing the currently NULL cfg_sink. E.g: ring x ring x # fails on "already exists" server srv 1.1.1.1 # crashes on cfg_sink==NULL All other statements have a test for this but "server" was missing it, so this patch adds it. Thanks to Joel Hutchinson for reporting this issue. This must be backported as far as 2.2.	2022-11-16 18:59:43 +01:00
Willy Tarreau	9fd0542148	MEDIUM: trace: create a new "trace" statement in the "global" section The exact same commands as those from the CLI may be pre-loaded at boot time by passing them one per line after the "trace" keyword in the global section; i.e. just copy-pasting all commands directly there will do the job. Note that if a ring is mentioned, it needs to be declared before the global section. Another option is to append another global section after "ring". For now the keyword is marked as experimental to discourage its broad adoption by default. "expose-experimental-directives" needs to be placed in the global section to expose it.	2022-11-16 17:55:53 +01:00
Willy Tarreau	c11f1cdf4d	MINOR: trace: split the CLI "trace" parser in CLI vs statement In order to be able to reuse the "trace" statements elsewhere (e.g. global section), we'll first need to split its parser. It turns out that the whole thing is self-contained inside a single function that emits a single message on warning/error or nothing on success. That's quite easy to split in two parts, the one that does the job and produces the status message and the one that sends it to the CLI. That's what this patch does.	2022-11-16 17:55:53 +01:00
Mickael Torres	226082d13a	BUG/MINOR: mux-h1: Do not send a last null chunk on body-less answers HEAD answers should not contain any body data. Currently when a "transfer-encoding: chunked" header is returned, a last null-chunk is added to the answer. Some clients choke on it and fail when trying to reuse the connection. Check that the response should not be body-less before sending the null-chunk. This patch should fix #1932. It must be backported as far as 2.4.	2022-11-16 16:25:26 +01:00
Remi Tricot-Le Breton	e608b0eb16	BUG/MINOR: ssl: SSL_load_error_strings might not be defined The SSL_load_error_strings function was marked as deprecated in OpenSSL 1.1.0 so compiling HAProxy with OPENSSL_NO_DEPRECATED set and a recent OpenSSL library would fail. The manpages say that this function was replaced by OPENSSL_init_crypto and OPENSSL_init_ssl which are already called at start up by the SSL lib. We do not seem to be in a case where explicit call of those functions is required. This patch fixes GitHub issue #1813. It can be backported to 2.6.	2022-11-16 11:09:33 +01:00
Christopher Faulet	52fd8a1b7b	BUG/MEDIUM: mux-fcgi: Avoid value length overflow when it doesn't fit at once When the request data are copied in a mbuf, if the free space is too small to copy all data at once, the data length is shortened. When this is performed, we reserve the size of the STDIN recod header and eventually the same for the empty STDIN record if it is the last HTX block of the request. However, there is no test to be sure the free space is large enough. Thus, on this special case, when the mbuf is almost full, it is possible to overflow the value length. Because of this bug, it is possible to experience crashes from time to time. This patch should fix the issue #1923. It must be backported as far as 2.4.	2022-11-16 09:27:09 +01:00
Christopher Faulet	e8c7fb3588	BUG/MINOR: mux-fcgi: Be sure to send empty STDING record in case of zero-copy When the last HTX DATA block was copied in zero-copy, the empty STDIN record, marking the end of the request data was never sent. Thanks to this patch, it is now sent. This patch must be backported as far as 2.4.	2022-11-16 09:27:09 +01:00
Christopher Faulet	2364b39984	BUG/MINOR: resolvers: Set port before IP address when processing SRV records For a server subject to SRV resolution, when the server's address is set, its dynamic cookie, if any, and its server key are computed. Both are based on the ip/port pair. However, this happens before the server's port is set. Thus the port is equal to 0 at this stage. It is a problem if several servers share the same IP but with different ports because they will share the same dynamic cookie and the same server key, disturbing this way the connection persistency and the session stickiness. This patch must be backported as far as 2.2.	2022-11-16 09:27:09 +01:00
Christopher Faulet	68a61b6321	BUG/MINOR: resolvers: Don't wait periodic resolution on healthcheck failure DNS resoltions may be triggered via a "do-resolve" action or when a connection failure is experienced during a healthcheck. Cached valid responses are used, if possible. But if the entry is expired or if there is no valid response, a new reolution should be performed. However, an resolution is only performed if the "resolve" timeout is expired. Thus, when this comes from a healthcheck, it means no extra resolution is performed at all. Now, when the resolution is performed for a server (SRV or SRVEQ) and no valid response is found, the resolution timer is reset (last_resolution is set to TICK_ETERNITY). Of course, it is only performed if no resolution is already running. Note that this feature was broken 5 years ago when the resolvers code was refactored (`67957bd59e`). This patch should fix the issue #1906. It affects all stable versions. However, it is probably a good idea to not backport it too far (2.6, maybe 2.4) and with some delay.	2022-11-16 09:27:09 +01:00
Christopher Faulet	5a3d9a77e2	BUG/MINOR: http-htx: Fix error handling during parsing http replies When an error is triggered during arguments parsing of an http reply (for instance, from a "return" rule), while a log-format body was expected but not evaluated yet, HAproxy crashes when the body log-format string is released because it was not properly initialized. The list used for the log-format string must be initialized earlier. This patch should fix the issue #1925. It must be backported as far as 2.2.	2022-11-16 09:27:09 +01:00
William Lallemand	78c7a06e4f	MINOR: ssl: reintroduce ERR_GET_LIB(ret) == ERR_LIB_PEM in ssl_sock_load_pem_into_ckch() Commit `432cd1a` ("MEDIUM: ssl: be stricter about chain error") removed the ERR_GET_LIB(ret) != ERR_LIB_PEM to be stricter about errors. However, PEM_R_NO_START_LINE is better be checked with ERR_LIB_PEM. So this patch complete the previous one. The original problem was that the condition was wrongly inversed. This original code from openssl: if (ERR_GET_LIB(err) == ERR_LIB_PEM && ERR_GET_REASON(err) == PEM_R_NO_START_LINE) became: if (ret && (ERR_GET_LIB(ret) != ERR_LIB_PEM && ERR_GET_REASON(ret) != PEM_R_NO_START_LINE)) instead of: if (ret && !(ERR_GET_LIB(ret) == ERR_LIB_PEM && ERR_GET_REASON(ret) == PEM_R_NO_START_LINE)) This must not be backported as it will break a lot of setup. That's too bad because a lot of errors are lost. Not marked as a bug because of the breakage it could cause on working setups.	2022-11-15 18:24:17 +01:00
William Lallemand	45fed2c7a6	MINOR: ssl: ssl_sock_load_cert_chain() display error strings Display error strings when SSL_CTX_use_certificate() or SSL_CTX_set1_chain() doesn't work.	2022-11-15 16:56:03 +01:00
Willy Tarreau	e98d385819	MINOR: deinit: add a "quick-exit" option to bypass the deinit step Once in a while we spot a bug in the deinit code that is complex, especially when it has to deal with incomplete initializations, and the ability to bypass this step has regularly been raised. In addition for fast-reloading setups it could theoretically save some time. Tests have shown that very large configs can barely save ~100-150ms by skipping the deinit step. However the ability not to crash if a bug is encountered can occasionally help. This patch adds an option to do exactly this. It's obviously not enabled by default and the documentation discourages from using it, but this might be useful in the future.	2022-11-15 09:37:09 +01:00
Aurelien DARRAGON	16d6c0cb09	BUG/MEDIUM: wdt/clock: properly handle early task hangs In `ae053b30` - BUG/MEDIUM: wdt: don't trigger the watchdog when p is unitialized: wdt is not triggering until prev_cpu_time is initialized to prevent unexpected process termination. Unfortunately this is not enough, some tasks could start immediately after process startup, and in such cases prev_cpu_time could be uninitialized, because prev_cpu_time is set after the polling loop while process_runnable_tasks() is executed before the polling loop. It happens to be the case with lua tasks registered using register_task function from lua script. Those tasks are registered in early init stage of haproxy and they are scheduled to run before the first polling loop, leading to prev_cpu_time being uninitialized (equals 0) on the thread when the task is first executed. Because of this, if such tasks get stuck right away (e.g: blocking IO) the watchdog won't behave as expected and the thread will remain stuck indefinitely. (polling loop for the thread won't run at all as the thread is already stuck) To solve this, we're now making sure that prev_cpu_time is first set before any tasks are processed on the thread. This is done by setting initial prev_cpu_time value directly in clock_init_thread_date() Thanks to Abhijeet Rastogi for reporting this unexpected behavior. It could be backported in every stable versions. (everywhere `ae053b30` is, because both are related)	2022-11-14 19:14:53 +01:00
Willy Tarreau	aa1909edf7	MEDIUM: http-ana: remove set-cookie2 support This has never really been implemented in clients nor servers. We wanted to drop it from 2.5 already but forgot, so let's do it now. The code was only minimally changed. It could possibly be slightly simplified but it would only be marginal, at the great risk of breaking something, thus let's keep it in its proven state instead. Tracked in github issue #1551.	2022-11-14 19:01:45 +01:00
Willy Tarreau	fbb934da90	BUG/MEDIUM: stick-table: fix a race condition when updating the expiration task Pierre Cheynier reported a rare crash that can affect stick-tables. When a entry is created, the stick-table's expiration date is updated. But if at exactly the same time the expiration task runs, it finishes by updating its expiration timer without any protection, which may collide with the call to task_queue() in another thread. In this case, it sometimes happens that the first test for TICK_ETERNITY in task_queue() passes, then the "expire" field is reset, then the BUG_ON() triggers, like below: FATAL: bug condition "task->expire == 0" matched at src/task.c:279 call trace(13): \| 0x649d86 [c6 04 25 01 00 00 00 00]: __task_queue+0xc6/0xce \| 0x596bef [eb 90 ba 03 00 00 00 be]: stktable_requeue_exp+0x1ef/0x258 \| 0x596c87 [48 83 bb 90 00 00 00 00]: stktable_touch_with_exp+0x27/0x312 \| 0x563698 [48 8b 4c 24 18 4c 8b 4c]: stream_process_counters+0x3a8/0x6a2 \| 0x569344 [49 8b 87 f8 00 00 00 48]: process_stream+0x3964/0x3b4f \| 0x64a80b [49 89 c7 e9 23 ff ff ff]: run_tasks_from_lists+0x3ab/0x566 \| 0x64ad66 [29 44 24 14 8b 7c 24 14]: process_runnable_tasks+0x396/0x71e \| 0x6184b2 [83 3d 47 b3 a6 00 01 0f]: run_poll_loop+0x92/0x4ff \| 0x618acf [48 8b 1d aa 20 7d 00 48]: main+0x1877ef \| 0x7fc7d6ec1e45 [64 48 89 04 25 30 06 00]: libpthread:+0x7e45 \| 0x7fc7d6c9e4af [48 89 c7 b8 3c 00 00 00]: libc:clone+0x3f/0x5a This one is extremely difficult to reproduce in practice, but adding a printf() in process_table_expire() before assigning the value, while running with an expire delay of 1ms helps a lot and may trigger the crash in less than one minute on a 8-thread machine. Interestingly, depending on the sequencing, this bug could also have made a table fail to expire if the expire field got reset after the last update but before the call to task_queue(). It would require to be quite unlucky so that the table is never touched anymore after the race though. The solution taken by this patch is to take the table's lock when updating its expire value in stktable_requeue_exp(), enclosing the call to task_queue(), and to update the task->expire field while still under the lock in process_table_expire(). Note that thanks to previous changes, taking the table's lock for the update in stktable_requeue_exp() costs almost nothing since we now have the guarantee that this is not done more than 1000 times a second. Since process_table_expire() sets the timeout after returning from stktable_trash_expired() which just released the lock, the two functions were merged so that the task's expire field is updated while still under the lock. Note that this heavily depends on the two previous patches below: CLEANUP: stick-table: remove the unused table->exp_next OPTIM: stick-table: avoid atomic ops in stktable_requeue_exp() when possible This is a bit complicated due to the fact that in 2.7 some parts were made lockless. In 2.6 and older, the second part (the merge of the two functions) will be sufficient since the task_queue() call was already performed under the table's lock, and the patches above are not needed. This needs to be backported as far as 1.8 scrupulously following instructions above.	2022-11-14 18:20:38 +01:00
Willy Tarreau	3238f79d12	OPTIM: stick-table: avoid atomic ops in stktable_requeue_exp() when possible Since the task's time resolution is the millisecond we know there will not be more than 1000 useful updates per second, so there's no point in doing a CAS and a task_queue() for each call, better first check if we're going to change the date. Now we're certain not to perform such operations more than 1000 times a second for a given table. The loop was modified because this improvement will also be used to fix a bug later.	2022-11-14 18:20:38 +01:00
Willy Tarreau	6342714052	CLEANUP: stick-table: remove the unused table->exp_next The ->exp_next field of the stick-table was probably useful in 1.5 but it currently only carries a copy of what the future value of the table's task's expire value will be, while it's systematically copied over there immediately after being assigned. As such it provides exactly a local variable. Let's remove it, as it costs atomic operations.	2022-11-14 18:20:38 +01:00
William Lallemand	f813dab175	BUG/MINOR: ssl: crt-ignore-err memory leak with 'all' parameter Only allocate "str" if the parameter is not "all" in order to avoid any memory leak. No backport needed.	2022-11-14 11:43:52 +01:00
William Lallemand	380085dfc1	CLEANUP: ssl: remove printf in bind_parse_ignore_err Remove debug printf. No backport needed.	2022-11-14 11:34:07 +01:00
Willy Tarreau	476c2802c5	BUILD: stconn: use __fallthrough in various shutw() functions This avoids 7 build warnings when preprocessing happens before compiling with gcc >= 7.	2022-11-14 11:14:02 +01:00
Willy Tarreau	91d398cce2	BUILD: compression: use __fallthrough in comp_http_payload() This avoids one build warning when preprocessing happens before compiling with gcc >= 7.	2022-11-14 11:14:02 +01:00
Willy Tarreau	ab42dc3358	BUILD: map: use __fallthrough in cli_io_handler_*() This avoids 6 build warnings when preprocessing happens before compiling with gcc >= 7.	2022-11-14 11:14:02 +01:00
Willy Tarreau	68a71ed3f6	BUILD: vars: use __fallthrough in var_accounting_{diff,add}() This avoids 6 build warnings when preprocessing happens before compiling with gcc >= 7.	2022-11-14 11:14:02 +01:00
Willy Tarreau	779aa693d9	BUILD: h1_htx: use __fallthrough in h1_parse_chunk() This avoids one build warning when preprocessing happens before compiling with gcc >= 7.	2022-11-14 11:14:02 +01:00
Willy Tarreau	c5bc4ad24d	BUILD: http_act: use __fallthrough in parse_http_del_header() This avoids one build warning when preprocessing happens before compiling with gcc >= 7.	2022-11-14 11:14:02 +01:00
Willy Tarreau	36a73439f9	BUILD: check: use __fallthrough in __health_adjust() This avoids one build warning when preprocessing happens before compiling with gcc >= 7.	2022-11-14 11:14:02 +01:00
Willy Tarreau	80f9a63184	BUILD: logs: use __fallthrough in build_log_header() This avoids 4 build warnings when preprocessing happens before compiling with gcc >= 7.	2022-11-14 11:14:02 +01:00
Willy Tarreau	3064f52b15	BUILD: spoe: use __fallthrough in spoe_handle_appctx() This avoids two build warnings when preprocessing happens before compiling with gcc >= 7.	2022-11-14 11:14:02 +01:00
Willy Tarreau	8de35935b0	BUILD: acl: use __fallthrough in parse_acl_expr() This avoids two build warnings when preprocessing happens before compiling with gcc >= 7.	2022-11-14 11:14:02 +01:00
Willy Tarreau	266ce55109	BUILD: args: use __fallthrough in make_arg_list() This avoids one build warning when preprocessing happens before compiling with gcc >= 7.	2022-11-14 11:14:02 +01:00
Willy Tarreau	7de8de0bf8	BUILD: tools: use __fallthrough in url_decode() This avoids one build warning when preprocessing happens before compiling with gcc >= 7.	2022-11-14 11:14:02 +01:00
Willy Tarreau	cff89874ea	BUILD: hash: use __fallthrough in hash_djb2() This avoids 5 build warnings when preprocessing happens before compiling with gcc >= 7.	2022-11-14 11:14:02 +01:00
Willy Tarreau	7910825409	BUILD: peers: use __fallthrough in peer_io_handler() This avoids 7 build warnings when preprocessing happens before compiling with gcc >= 7.	2022-11-14 11:14:02 +01:00
Willy Tarreau	9eb93c031b	BUILD: stats: use __fallthrough in stats_dump_proxy_to_buffer() This avoids 12 build warnings when preprocessing happens before compiling with gcc >= 7.	2022-11-14 11:14:02 +01:00
Willy Tarreau	f3f60763fa	BUILD: tcpcheck: use __fallthrough in check_proxy_tcpcheck() This avoids 1 build warning when preprocessing happens before compiling with gcc >= 7.	2022-11-14 11:14:02 +01:00
Willy Tarreau	469847945c	BUILD: stream: use __fallthrough in stats_dump_full_strm_to_buffer() This avoids 1 build warning when preprocessing happens before compiling with gcc >= 7.	2022-11-14 11:14:02 +01:00
Willy Tarreau	14de395a30	BUILD: hlua: use __fallthrough in hlua_post_init_state() This avoids 5 build warnings when preprocessing happens before compiling with gcc >= 7.	2022-11-14 11:14:02 +01:00
Willy Tarreau	a551f4fbfd	BUILD: ssl: use __fallthrough in cli_io_handler_tlskeys_files() This avoids two build warnings when preprocessing happens before compiling with gcc >= 7.	2022-11-14 11:14:02 +01:00
Willy Tarreau	6fcc86b312	BUILD: ssl: use __fallthrough in cli_io_handler_commit_{cert,cafile_crlfile}() This avoids 8 build warnings when preprocessing happens before compiling with gcc >= 7.	2022-11-14 11:14:02 +01:00
Willy Tarreau	aef8448b58	BUILD: ssl/crt-list: use __fallthrough in cli_io_handler_add_crtlist() This avoids 3 build warnings when preprocessing happens before compiling with gcc >= 7.	2022-11-14 11:14:02 +01:00
Willy Tarreau	5c8b52f80a	BUILD: quic: use __fallthrough in quic_connect_server() This avoids one build warning when preprocessing happens before compiling with gcc >= 7.	2022-11-14 11:14:02 +01:00
Willy Tarreau	7ed0597ce8	BUILD: sample: use __fallthrough in smp_is_rw() and smp_dup() This avoids three build warnings when preprocessing happens before compiling with gcc >= 7.	2022-11-14 11:14:02 +01:00
Willy Tarreau	eda36f1c23	IMPORT: slz: declare len to fix debug build when optimal match is enabled Building with -DFIND_OPTIMAL_MATCH would fail on undeclared "len". This one likely vanished in some cleanup. This is libslz upstream commit 1ea20360715e1ad0cd81db83fa4361310716b8cc	2022-11-14 11:14:02 +01:00
Willy Tarreau	a8a83bcc80	BUILD: ssl_utils: fix build on gcc versions before 8 Commit `960fb74ca` ("MEDIUM: ssl: {ca,crt}-ignore-err can now use error constant name") provided a very convenient way to initialize only desired macros. Unfortunately with gcc versions older than 8, it breaks with: src/ssl_utils.c:473:12: error: initializer element is not constant because it seems that the compiler cannot resolve strings to constants at build time. This patch takes a different approach, it stores the value of the macro as a string and this string is converted to integer at boot time. This way it works everywhere.	2022-11-14 11:12:49 +01:00
William Lallemand	4639689d89	BUG/MINOR: ssl: bind_conf is uncorrectly accessed when using QUIC Since commit 9b2598 ("BUG/MEDIUM: ssl: Verify error codes can exceed 63"), the ca_ignerr_bitfield and crt_ignerr_bietfield are incorrecly accessed from __objt_listener(conn->target)->bind_conf which is not avaiable from QUIC. The bind_conf variable was mistakenly replaced. This patch fixes the issue by using again the bind_conf variable. Must be backported where 9b2598 was backported.	2022-11-10 16:56:21 +01:00
Amaury Denoyelle	30fc6da148	MINOR: server: clear prefix on stderr logs after add server cli_parse_add_server() is the CLI handler for 'add server' command. This functions uses usermsgs_ctx to retrieve logs messages from internal ha_alert() calls and display it at the end of the handler. At the beginning of the handler, stderr prefix is defined to "CLI" via usermsgs_clr() function. However, this is not resetted at the end. This causes inconsistency for stderr output : 1. each ha_alert() invocation will reuse "CLI" prefix if 'add server' command was executed before, even in non-CLI context 2. usermsgs_ctx is thread local, so this is only true if this runs on the same thread as 'add server' handler. To fix this, ensure that "CLI" prefix is now resetted after cli_parse_add_server(). This is done thanks to the addition to cli_umsg()/cli_umsgerr() functions. This can be backported up to 2.5 if we prefer to ensure output consistency at the risk of changing stderr behaviors in stable versions. In this case, the previous commit should be backported before : MINOR: cli: define usermsgs print context	2022-11-10 16:42:47 +01:00
Amaury Denoyelle	24e9961a8f	MINOR: cli: define usermsgs print context CLI 'add server' handler relies on usermsgs_ctx to display errors in internal function on CLI output. This may be also extended to other handlers. However, to not clutter stderr from another contextes, usermsgs_ctx must be resetted when it is not needed anymore. This operation cannot be conducted in the CLI parse handler as display is conducted after it. To achieve this, define new CLI states CLI_ST_PRINT_UMSG / CLI_ST_PRINT_UMSGERR. Their principles is nearly identical to states for dynamic messages printing.	2022-11-10 16:42:47 +01:00
Amaury Denoyelle	56f50a03b7	CLEANUP: cli: rename dynamic error printing state Rename CLI_ST_PRINT_FREE to CLI_ST_PRINT_DYNERR. Most notably, this highlights that this is reserved to error printing. This is done to ensure consistency between CLI_ST_PRINT/CLI_ST_PRINT_DYN and CLI_ST_PRINT_ERR/CLI_ST_PRINT_DYNERR. The name is also consistent with the function cli_dynerr() which activates it.	2022-11-10 16:42:47 +01:00
William Lallemand	9fbc84e571	MINOR: ssl: x509_v_err_str converter transforms an integer to a X509_V_ERR name The x509_v_err_str converter transforms a numerical X509 verify error to its constant name.	2022-11-10 13:28:37 +01:00
William Lallemand	960fb74cae	MEDIUM: ssl: {ca,crt}-ignore-err can now use error constant name The ca-ignore-err and crt-ignore-err directives are now able to use the openssl X509_V_ERR constant names instead of the numerical values. This allow a configuration to survive an OpenSSL upgrade, because the numerical ID can change between versions. For example X509_V_ERR_INVALID_CA was 24 in OpenSSL 1 and is 79 in OpenSSL 3. The list of errors must be updated when a new major OpenSSL version is released.	2022-11-10 13:28:37 +01:00
Remi Tricot-Le Breton	9b25982716	BUG/MEDIUM: ssl: Verify error codes can exceed 63 The CRT and CA verify error codes were stored in 6 bits each in the xprt_st field of the ssl_sock_ctx meaning that only error code up to 63 could be stored. Likewise, the ca-ignore-err and crt-ignore-err options relied on two unsigned long longs that were used as bitfields for all the ignored error codes. On the latest OpenSSL1.1.1 and with OpenSSLv3 and newer, verify errors have exceeded this value so these two storages must be increased. The error codes will now be stored on 7 bits each and the ignore-err bitfields are replaced by a big enough array and dedicated bit get and set functions. It can be backported on all stable branches. [wla: let it be tested a little while before backport] Signed-off-by: William Lallemand <wlallemand@haproxy.org>	2022-11-10 11:45:48 +01:00
Remi Tricot-Le Breton	aa529f776d	BUG/MINOR: ssl: ocsp structure not freed properly in case of error In case of error, the ocsp item might already be in the ocsp certificate tree but simply freed instead of destroyed through ssl_sock_free_ocsp. This patch can be backported to all stable versions.	2022-11-04 11:40:29 +01:00
Remi Tricot-Le Breton	1621dc1cc5	BUG/MINOR: ssl: Memory leak of AUTHORITY_KEYID struct when loading issuer When calling ssl_get0_issuer_chain, if akid is not NULL but its keyid is, then the AUTHORITY_KEYID is not freed. This patch can be backported to all stable branches.	2022-11-04 11:40:29 +01:00
Remi Tricot-Le Breton	a2c21db155	BUG/MINOR: ssl: Memory leak of DH BIGNUM fields When running HAProxy with OpenSSLv3, the two BIGNUMs used to build our own DH parameters are not freed. It was not necessary previously because ownership of those parameters was transferred to OpenSSL through the DH_set0_pqg call. This patch should be backported to 2.6.	2022-11-04 11:40:29 +01:00
Miroslav Zagorac	a2ec192de3	BUG/MINOR: httpclient: fixed memory allocation for the SSL ca_file The memory for the SSL ca_file was allocated only once (in the function httpclient_create_proxy()) and that pointer was assigned to each created proxy that the HTTP client uses. This would not be a problem if this memory was not freed in each individual proxy when it was deinitialized in the function ssl_sock_free_srv_ctx(). Memory allocation: src/http_client.c, function httpclient_create_proxy(): 1277: if (!httpclient_ssl_ca_file) 1278: httpclient_ssl_ca_file = strdup("@system-ca"); 1280: srv_ssl->ssl_ctx.ca_file = httpclient_ssl_ca_file; Memory deallocation: src/ssl_sock.c, function ssl_sock_free_srv_ctx(): 5613: ha_free(&srv->ssl_ctx.ca_file); This should be backported to version 2.6.	2022-11-04 11:29:18 +01:00
William Lallemand	1ef1b859d0	CLEANUP: ssl: remove dead code in ssl_sock_load_pem_into_ckch() Commit `432cd1a` ("MEDIUM: ssl: be stricter about chain error") introduced some dead code, let's remove it. Should fix issue #1909.	2022-10-30 19:00:06 +01:00
Ilya Shipitsin	4a689dad03	CLEANUP: assorted typo fixes in the code and comments This is 32nd iteration of typo fixes	2022-10-30 17:17:56 +01:00
Amaury Denoyelle	0b13e94071	BUG/MINOR: quic: fix race condition on datagram purging Each datagram is received by a random thread and dispatch to its destination thread linked to the connection. Then, the datagram is handled by the connection thread. Once this is done, datagram buffer pointer is atomically set to NULL to mark it as consumed. Consumed datagrams are purged before recvfrom() invocation on random receiver threads. The check for NULL buffer must thus be done atomically. This was not the case before this patch, which may have triggered race conditions. This bug has been introduced by commit `91b2305ad7` MINOR: quic: implement datagram cleanup for quic_receiver_buf This should be backported up to 2.6 after previously mentionned commit.	2022-10-27 18:35:49 +02:00
Amaury Denoyelle	735b44f5df	MINOR: quic: add counter for interrupted reception Add a new counter "quic_rxbuf_full". It is incremented each time quic_sock_fd_iocb() is interrupted on full buffer. This should help to debug github issue #1903. It is suspected that QUIC receiver buffers are full which in turn cause quic_sock_fd_iocb() to be called repeatedly resulting in a high CPU consumption.	2022-10-27 18:35:42 +02:00
William Lallemand	5de4951252	MINOR: ssl: dump the SSL string error when SSL_CTX_use_PrivateKey() failed. Display the OpenSSL reason error string when SSL_CTX_use_PrivateKey() failed.	2022-10-27 14:50:22 +02:00
Aurelien DARRAGON	7faffdc6ab	BUG/MINOR: log: fixing bug in tcp syslog_io_handler Octet-Counting syslog_io_handler does specific treatment to handle syslog tcp octet counting: Logic was good, but a sneaky mistake prevented rfc-6587 octet counting from working properly. trash.area was used as an input buffer. It does not make sense here since it is uninitialized. Compilation was unaffected because trash is a thread local "global" variable. buf->area should definitely be used instead. This should be backported as far as 2.4.	2022-10-27 11:28:53 +02:00
Amaury Denoyelle	bbb1c68508	BUG/MINOR: quic: fix subscribe operation Subscribing was not properly designed between quic-conn and quic MUX layers. Align this as with in other haproxy components : <subs> field is moved from the MUX to the quic-conn structure. All mention of qcc MUX is cleaned up in quic_conn_subscribe()/quic_conn_unsubscribe(). Thanks to this change, ACK reception notification has been simplified. It's now unnecessary to check for the MUX existence before waking it. Instead, if <subs> quic-conn field is set, just wake-up the upper layer tasklet without mentionning MUX. This should probably be extended to other part in quic-conn code. This should be backported up to 2.6.	2022-10-26 18:18:26 +02:00
Amaury Denoyelle	0aba11e9e7	MINOR: quic: remove unnecessary quic_session_accept() A specialized listener accept was previously used for QUIC. This is now unneeded and we can revert to the default one session_accept_fd(). One change of importance is that the call order between conn_xprt_start() and conn_complete_session() is now reverted to the default one. This means that MUX instance is now NULL during qc_xprt_start() and its app-ops layer cannot be set here. This operation has been delayed to qc_init() to prevent a segfault. This should be backported up to 2.6.	2022-10-26 18:16:20 +02:00
Christopher Faulet	b976640fe1	BUG/MAJOR: stick-table: don't process store-response rules for applets The commit `bc7c207f74` ("BUG/MAJOR: stick-tables: do not try to index a server name for applets") tried to catch applets case when we tried to index the server name. However, there is still an issue. The applets are unconditionally casted to servers and this bug exists since a while. it's just luck if it doesn't crash. Now, when store rules are processed, we skip the rule if the stream's target is not a server or, of course, if it is a server but the "non-stick" option is set. However, we still take care to release the sticky session. This patch must be backported to all stable versions.	2022-10-25 18:04:54 +02:00
William Lallemand	432cd1a7f8	MEDIUM: ssl: be stricter about chain error The error check on certificate chain was ignoring all decoding error, silently ignoring some errors. This patch fixes the issue by being stricter on errors when reading the chain, this is a change of behavior, it could break existing setup that has a wrong chain.	2022-10-25 15:55:13 +02:00
William Lallemand	a538452fa4	MINOR: ssl: add the SSL error string before the chain Add the SSL error string when failing to load a certificate in ssl_sock_load_pem_into_ckch(). It's difficult to know what happen when no descriptive errror are emitted. This one is for the certificate before trying to load the complete chain.	2022-10-25 15:53:01 +02:00
William Lallemand	f784b90eae	MINOR: ssl: add the SSL error string when failing to load a certificate Add the SSL error string when failing to load a certificate in ssl_sock_load_pem_into_ckch(). It's difficult to know what happen when no descriptive errror are emitted. Example: [ALERT] (1264006) : config : parsing [ssl_default_server.cfg:51] : 'bind /tmp/ssl.sock' in section 'listen' : unable to load certificate chain from file 'reg-tests/ssl//common.pem': ASN no PEM Header Error	2022-10-25 12:36:48 +02:00
Christopher Faulet	d08a25b1f1	BUG/MINOR: sink: Set default connect/server timeout for implicit ring buffers Ring buffers may be implicitly created from log declarations when "tcp@", "tcp6@", "tcp4@" or "uxst@" prefixes are used. These ring buffers rely on unconfigurable proxies. While connect and server timeouts should be defined for explicit ring buffers, it is no possible for implicit ones. However, a default value must be set and TICK_ETERNITY is not an acceptable one. Thus, now "1s" is set for the connect timeout and "5s" is set for server one. This patch may be backported as far as 2.4.	2022-10-24 16:00:49 +02:00
Christopher Faulet	11a707ae52	BUG/MINOR: sink: Only use backend capability for the sink proxies When a ring section is parsed, a proxy is created. For now, it has the frontend (PR_CAP_FE) and the internal (PR_CAP_INT) capabilities, in addition to the expected backend capability (PR_CAP_BE). PR_CAP_INT capability was added to silent warning triggered because of PR_CAP_FE capability. Indeed, Because the proxy is declared as a frontend, warnings about missing bind lines and missing client timeout should be triggered during the configuration parsing. These warnings are inhibited because PR_CAP_INT capability is set. It is an issue on the 2.4 because PR_CAP_INT capability does not exist. So warnings are always emitted. But the true bug is that these proxies should not have PR_CAP_FE and PR_CAP_INT capabilities. Removing these capabilities is enough to remove any warnings on the 2.4, with no regression on higher versions. However, it may be a good idea to eval if a dedicated frontend for sinks should be added or not. This way, a true frontend would be used to start the sink applets. In addition, proxies capabilities/modes have to be reviewed to have a less ambiguous API. For instance a dedicate mode for sinks (PR_MODE_SINK ?) may be added. Finally, it could be very nice to have all proxies in the same list, including internal ones. This patch should fix the issue #1900. It must be backported as far as 2.4.	2022-10-24 16:00:49 +02:00
Emeric Brun	ac556082e7	MINOR: peers: handle multiple resync requests using shards We considered the resync process is finished if a full resync request is ended receiving the "resync-finish" message. But in the case of "shards" each node declared with a "shard" has only a partial view of the table. And the resync process is ended whereas the original peer tables content contains only a "shard" of the full content. This patch allow to retrieve the entire tables requesting a resync from all different "shards". To do so we don't commit the end of a resync process receiving a "resync-finish" if the node is part of "shard", we only flag this peer and all peers using the same shard as "notup2date" as if we received a "resync-partial" message, and we re-schedule a request of a resync as it is done receiving a "resync-partial" message. Doing this the peers flagged "notup2date" won't be addressed for the next resync request round and the next resync request will be send to a shard not yet requested. Receving a "resync-finish" message we also check if all peers using "shards" are flagged "notup2date". It meens that all peers have been addressed and we can considered the resync process is now finished. Note also that the "resync request" scheduler already handle a timeout and if we are not able to retrieve a full resync after a delay. The resync process is ended. This patch should be backported in all versions handling "shard" on peer lines.	2022-10-24 10:55:53 +02:00
Fr�d�ric L�caille	36d1565640	MINOR: peers: Support for peer shards Add "shards" new keyword for "peers" section to configure the number of peer shards attached to such secions. This impact all the stick-tables attached to the section. Add "shard" new "server" parameter to configure the peers which participate to all the stick-tables contents distribution. Each peer receive the stick-tables updates only for keys with this shard value as distribution hash. The "shard" value is stored in ->shard new server struct member. cfg_parse_peers() which is the function which is called to parse all the lines of a "peers" section is modified to parse the "shards" parameter stored in ->nb_shards new peers struct member. Add srv_parse_shard() new callback into server.c to pare the "shard" parameter. Implement stksess_getkey_hash() to compute the distribution hash for a stick-table key as the 64-bits xxhash of the key concatenated to the stick-table name. This function is called by stksess_setkey_shard(), itself called by the already implemented function which create a new stick-table key (stksess_new()). Add ->idlen new stktable struct member to store the stick-table name length to not have to compute it each time a stick-table key hash is computed.	2022-10-24 10:55:53 +02:00
Amaury Denoyelle	7941ead3aa	MINOR: quic: display unknown error sendto counter on stat page This patch complete the previous incomplete commit. The new counter sendto_err_unknown is now displayed on stats page/CLI show stats. This is related to github issue #1903. This should be backported up to 2.6.	2022-10-24 10:52:59 +02:00
Amaury Denoyelle	1d9f170edd	MINOR: quic: do not crash on unhandled sendto error Remove ABORT_NOW() statement on unhandled sendto error. Instead use a dedicated counter sendto_err_unknown to report these cases. If we detect increment of this counter, strace can be used to detect errno value : $ strace -p $(pidof haproxy) -f -e trace=sendto -Z This should be backported up to 2.6. This should help to debug github issue #1903.	2022-10-24 10:18:44 +02:00
Christopher Faulet	910b7577bc	BUG/MEDIUM: compression: handle rewrite errors when updating response headers When an HTTP response is compressed by HAProxy, the headers are updated. However it is possible to encounter a rewrite error because the buffer is full. In this case, the compression is aborted. Thus, we must be sure to leave the response in a valid state. For now, it is an issue because the "Content-Encoding" header is added before all other headers manipulations. So if the compression is aborted on error, the "Content-Encoding" header may remain while the payload is not compressed. So now, we take care to leave with a valid response on error by reordering the headers manipulations. It is too painful to really rollback all changes, especially for an edge case. This patch should be backported as far as 2.0. Note that on the 2.0, the legacy HTTP part is also concerned.	2022-10-24 09:00:14 +02:00
Amaury Denoyelle	176174f7e4	BUG/MINOR: mux-quic: complete flow-control for uni streams Max stream data was not enforced and respect for local/remote uni streams. Previously, qcs instances incorrectly reused the limit defined from bidirectional ones. This is now fixed. Two fields are added in qcc structure connection : * value for local flow control to enforce on remote uni streams * value for remote flow control to respect on local uni streams These two values can be reused to properly initialized msd field of a qcs instance in qcs_new(). The rest of the code is similar. This must be backported up to 2.6.	2022-10-21 17:31:18 +02:00
William Lallemand	1344ebd74e	MINOR: mworker/cli: does no try to dump the startup-logs w/o USE_SHM_OPEN When haproxy is compiled without USE_SHM_OPEN, does not try to dump the startup-logs in the "reload" output, because it won't show anything interesting.	2022-10-21 14:03:29 +02:00
William Lallemand	2f67dd96ae	CLEANUP: mworker/cli: rename the status function to loadstatus clarify the name of the IO handler which show the reload status.	2022-10-21 14:03:06 +02:00
William Lallemand	a93eac41f0	BUG/MEDIUM: httpclient: check if the httpclient was released in the IO handler Upon a applet_release(), the applet can be scheduled again and a call to the IO handler is still possible. When the struct httpclient is already free the IO handler could try to access it. This patch fixes the issue by setting svcctx to NULL in the applet_release, and checking its value in the IO handler. Must be backported as far as 2.5.	2022-10-20 18:47:15 +02:00
William Lallemand	bb581423b3	BUG/MEDIUM: httpclient/lua: crash when the lua task timeout before the httpclient When the lua task finished before the httpclient that are associated to it, there is a risk that the httpclient try to task_wakeup() the lua task which does not exist anymore. To fix this issue the httpclient used in a lua task are stored in a list, and the httpclient are destroyed at the end of the lua task. Must be backported in 2.5 and 2.6.	2022-10-20 18:47:15 +02:00
Christopher Faulet	321d100cc8	BUG/MINOR: ring: Properly parse connect timeout The connect timeout in a ring section was not properly parsed. Thus, it was never set and the server timeout may be overwritten, depending on the directives order. The first char of the keyword must be tested, not the third one. This patch is related to the issue #1900. But it does not fix the issue. It must be backported as far as 2.4.	2022-10-20 09:03:19 +02:00
Christopher Faulet	cc640e851a	BUG/MINOR: log: Preserve message facility when the log target is a ring buffer When a ring is used as log target, the original facility, if any, must be preserved. The default facility must only be used if there no facility was found in the incoming log message. This patch should fix the issue #1901. It must be backported as far as 2.4.	2022-10-20 09:03:19 +02:00
Amaury Denoyelle	9e3026c58d	MINOR: quic: extend Retry token check function On Initial packet reception, token is checked for validity through quic_retry_token_check() function. However, some related parts were left in the parent function quic_rx_pkt_retrieve_conn(). Move this code directly into quic_retry_token_check() to facilitate its call in various context. The API of quic_retry_token_check() has also been refactored. Instead of working on a plain char* buffer, it now uses a quic_rx_packet instance. This helps to reduce the number of parameters. This change will allow to check Retry token even if data were received with a FD-owned quic-conn socket. Indeed, in this case, quic_rx_pkt_retrieve_conn() call will probably be skipped. This should be backported up to 2.6.	2022-10-19 18:45:58 +02:00
Amaury Denoyelle	6e56a9e055	MINOR: quic: refactor packet drop on reception Sometimes, a packet is dropped on reception. Several goto statements are used, mostly to increment a proxy drop counter or drop silently the packet. However, this labels are interleaved. Re-arrang goto labels to simplify this process : * drop label is used to drop a packet with counter incrementation. This is the default method. * drop_silent is the next label which does the same thing but skip the counter incrementation. This is useful when we do not need to report the packet dropping operation. This should be backported up to 2.6.	2022-10-19 18:45:58 +02:00
Amaury Denoyelle	982896961c	MINOR: quic: split and rename qc_lstnr_pkt_rcv() This change is the following of qc_lstnr_pkt_rcv() refactoring. This function has finally been split into several ones. The first half is renamed quic_rx_pkt_parse(). This function is responsible to parse a QUIC packet header and calculate the packet length. QUIC connection retrieval has been extracted and is now called directly by quic_lstnr_dghdlr(). The second half of qc_lstnr_pkt_rcv() is renamed to qc_rx_pkt_handle(). This function is responsible to copy a QUIC packet content to a quic-conn receive buffer. A third function named qc_rx_check_closing() is responsible to detect if the connection is already in closing state. As this requires to drop the whole datagram, it seems justified to be in a separate function. This change has no functional impact. It is part of a refactoring series on qc_lstnr_pkt_rcv(). The objective is to facilitate the integration of FD-owned quic-conn socket patches. This should be backported up to 2.6.	2022-10-19 18:45:55 +02:00
Amaury Denoyelle	449b1a8f55	MINOR: quic: extract connection retrieval Simplify qc_lstnr_pkt_rcv() by extracting code responsible to retrieve the quic-conn instance. This code is put in a dedicated function named quic_rx_pkt_retrieve_conn(). This new function could be skipped if a FD-owned quic-conn socket is used. The first traces of qc_lstnr_pkt_rcv() have been clean up as qc instance is always NULL here : thus qc parameter can be removed without any change. This change has no functional impact. It is a part of a refactoring series on qc_lstnr_pkt_rcv(). The objective is facilitate integration of FD-owned socket patches. This should be backported up to 2.6.	2022-10-19 18:12:56 +02:00
Amaury Denoyelle	deb7c87f55	MINOR: quic: define first packet flag Received packets treatment has some difference regarding if this is the first one or not of the encapsulating datagram. Previously, this was set via a function argument. Simplify this by defining a new Rx packet flag named QUIC_FL_RX_PACKET_DGRAM_FIRST. This change does not have functional impact. It will simplify API when qc_lstnr_pkt_rcv() is broken into several functions : their number of arguments will be reduced thanks to this patch. This should be backported up to 2.6.	2022-10-19 18:12:56 +02:00
Amaury Denoyelle	845169da58	MINOR: quic: extend pn_offset field from quic_rx_packet pn_offset field was only set if header protection cannot be removed. Extend the usage of this field : it is now set everytime on packet parsing in qc_lstnr_pkt_rcv(). This change helps to clean up API of Rx functions by removing unnecessary variables and function argument. This change has no functional impact. It is a part of a refactoring series on qc_lstnr_pkt_rcv(). The objective is facilitate integration of FD-owned socket patches. This should be backported up to 2.6.	2022-10-19 18:12:56 +02:00
Amaury Denoyelle	0eae57273b	MINOR: quic: add version field on quic_rx_packet Add a new field version on quic_rx_packet structure. This is set on header parsing in qc_lstnr_pkt_rcv() function. This change has no functional impact. It is a part of a refactoring series on qc_lstnr_pkt_rcv(). The objective is facilitate integration of FD-owned socket patches. This should be backported up to 2.6.	2022-10-19 18:12:56 +02:00
Amaury Denoyelle	6c940569f6	BUG/MINOR: quic: fix buffer overflow on retry token generation When generating a Retry token, client CID is used as encryption input. The client must reuse the same CID when emitting the token in a new Initial packet. A memory overflow can occur on quic_generate_retry_token() depending on the size of client CID. This is because space reserved for <aad> only accounted for QUIC_HAP_CID_LEN (size of haproxy owned generated CID). However, the client CID size only depends on client parameter and is instead limited to QUIC_CID_MAXLEN as specified in RFC9000. This was reproduced with ngtcp2 and haproxy built with ASAN. Here is the error log : ==14964==ERROR: AddressSanitizer: stack-buffer-overflow on address 0x7fffee228cee at pc 0x7ffff785f427 bp 0x7fffee2289e0 sp 0x7fffee228188 WRITE of size 17 at 0x7fffee228cee thread T5 #0 0x7ffff785f426 in __interceptor_memcpy /usr/src/debug/gcc/libsanitizer/sanitizer_common/sanitizer_common_interceptors.inc:827 #1 0x555555906ea7 in quic_generate_retry_token_aad src/quic_conn.c:5452 #2 0x555555907e72 in quic_retry_token_check src/quic_conn.c:5577 #3 0x55555590d01e in qc_lstnr_pkt_rcv src/quic_conn.c:6103 #4 0x5555559190fa in quic_lstnr_dghdlr src/quic_conn.c:7179 #5 0x555555eb0abf in run_tasks_from_lists src/task.c:590 #6 0x555555eb285f in process_runnable_tasks src/task.c:855 #7 0x555555d9118f in run_poll_loop src/haproxy.c:2853 #8 0x555555d91f88 in run_thread_poll_loop src/haproxy.c:3042 #9 0x7ffff709f8fc (/usr/lib/libc.so.6+0x868fc) #10 0x7ffff7121a5f (/usr/lib/libc.so.6+0x108a5f) This must be backported up to 2.6.	2022-10-18 14:36:47 +02:00
Frédéric Lécaille	ea492e3e47	BUILD: quic: Fix build for m68k cross-compilation Fix several warinings as this one: src/qmux_trace.c:80:45: error: format ‘%lu’ expects argument of type ‘long unsigned int’, but argument 4 has type ‘uint64_t’ {aka ‘const long long unsigned int’} [-Werror=format=] 80 \| chunk_appendf(&trace_buf, " qcs=%p .id=%lu .st=%s", \| ~~^ \| \| \| long unsigned int \| %llu 81 \| qcs, qcs->id, \| ~~~~~~~ \| \| \| uint64_t {aka const long long unsigned int} compilation terminated due to -Wfatal-errors. Cast remaining uint64_t variables as ullong with %llu as printf format and size_t others as ulong with %lu as printf format. Thank you to Ilya for having reported this issue in GH #1899. Must be backported to 2.6	2022-10-18 12:04:10 +02:00
Amaury Denoyelle	ba303deadc	BUILD: ssl_sock: fix null dereference for QUIC build A previous commit tries to fix uninitialized GCC warning on ssl code for QUIC build. See the fix here : `48e46f98cc` BUILD: ssl_sock: bind_conf uninitialized in ssl_sock_bind_verifycbk() However, this is incomplete as it still reports possible NULL dereference on ctx variable (GCC v12.2.0). Here is the compilation result : src/ssl_sock.c: In function ‘ssl_sock_bind_verifycbk’: src/ssl_sock.c:1739:12: error: potential null pointer dereference [-Werror=null-dereference] 1739 \| ctx->xprt_st \|= SSL_SOCK_ST_FL_VERIFY_DONE; \| To fix this, remove check on qc which can also never happens and replace it with a BUG_ON. This seems to satisfy GCC on my machine. This must be backported up to 2.6.	2022-10-17 18:58:09 +02:00
Thierry Fournier	74a9eb5216	BUG/MEDIUM: httpclient: segfault when the httpclient parser fails If the uri is unexpected ("/" in place of "http://xxx/"), some parsing function fails. The failure is not handled. This patch handle these errors. Note: the return code is boolean, maybe we can return more precise error for Lua reporting ? Must be backported in 2.6.	2022-10-17 12:04:06 +02:00
Fr�d�ric L�caille	5a5d05c71b	BUILD: quic: QUIC mux build fix for 32-bit build Thank you to Ilya for having reported this issue in GH #1897 Must be backported to 2.6.	2022-10-14 22:43:08 +02:00
Christopher Faulet	380ae9c3ff	MINOR: httpclient/lua: Don't set req_payload callback if body is empty The HTTPclient callback req_payload callback is set when a request payload must be streamed. In the lua, this callback is set when a body is passed as argument in one of httpclient functions (head/get/post/put/delete). However, there is no reason to set it if body string is empty. This patch is related to the issue #1898. It may be backported as far as 2.5.	2022-10-14 15:18:25 +02:00
Christopher Faulet	48005de17c	BUG/MEDIUM: httpclient: Don't set EOM flag on an empty HTX message In the HTTP client, when the request body is streamed, at the end of the payload, we must be sure to not set the EOM flag on an empty message. Otherwise, because there is no data, the buffer is reset to be released and the flag is lost. Thus, the HTTP client is never notified of the end of payload for the request and the applet is blocked. If the HTTP client is instanciated from a Lua script, it is even worse because we fall into a wakeup loop between the lua script and the HTTP client applet. At the end, HAProxy is killed because of the watchdog. This patch should fix the issue #1898. It must be backported to 2.6.	2022-10-14 15:18:25 +02:00
Fr�d�ric L�caille	48e46f98cc	BUILD: ssl_sock: bind_conf uninitialized in ssl_sock_bind_verifycbk() Even if this cannot happen, ensure <bind_conf> is initialized in this function to please some compilers. Takes the opportunity of this patch to replace an ABORT_NOW() by a BUG_ON() because if the variable values they test are not initialized, this is really because there is a bug. Must be backported to 2.6.	2022-10-14 10:25:11 +02:00
Willy Tarreau	f5a0c8abf5	MEDIUM: quic: respect the threads assigned to a bind line Right now the QUIC thread mapping derives the thread ID from the CID by dividing by global.nbthread. This is a problem because this makes QUIC work on all threads and ignores the "thread" directive on the bind lines. In addition, only 8 bits are used, which is no more compatible with the up to 4096 threads we may have in a configuration. Let's modify it this way: - the CID now dedicates 12 bits to the thread ID - on output we continue to place the TID directly there. - on input, the value is extracted. If it corresponds to a valid thread number of the bind_conf, it's used as-is. - otherwise it's used as a rank within the current bind_conf's thread mask so that in the end we still get a valid thread ID for this bind_conf. The extraction function now requires a bind_conf in order to get the group and thread mask. It was better to use bind_confs now as the goal is to make them support multiple listeners sooner or later.	2022-10-13 18:08:05 +02:00
William Lallemand	ec1f8a62ca	MINOR: mworker/cli: reload command displays the startup-logs Change the output of the "reload" command, it now displays "Success=0" if the reload failed and "Success=1" if it succeed. If the startup-logs is available (USE_SHM_OPEN=1), the command will print a "--\n" line, followed by the content of the startup-logs. Example: $ echo "reload" \| socat /tmp/master.sock - Success=1 -- [NOTICE] (482713) : haproxy version is 2.7-dev7-4827fb-69 [NOTICE] (482713) : path to executable is ./haproxy [WARNING] (482713) : config : 'http-request' rules ignored for proxy 'frt1' as they require HTTP mode. [NOTICE] (482713) : New worker (482720) forked [NOTICE] (482713) : Loading success. $ echo "reload" \| socat /tmp/master.sock - Success=0 -- [NOTICE] (482886) : haproxy version is 2.7-dev7-4827fb-69 [NOTICE] (482886) : path to executable is ./haproxy [ALERT] (482886) : config : parsing [test3.cfg:1]: unknown keyword 'Aglobal' out of section. [ALERT] (482886) : config : Fatal errors found in configuration. [WARNING] (482886) : Loading failure! $	2022-10-13 17:59:48 +02:00
William Lallemand	eba6a54cd4	MINOR: logs: startup-logs can use a shm for logging the reload When compiled with USE_SHM_OPEN=1 the startup-logs are now able to use an shm which is used to keep the logs when switching to mworker wait mode. This allows to keep the failed reload logs. When allocating the startup-logs at first start of the process, haproxy will do a shm_open with a unique path using the PID of the process, the file is unlink immediatly so we don't let unwelcomed files be. The fd resulting from this shm is stored in the HAPROXY_STARTUPLOGS_FD environment variable so it can be mmap again when switching to wait mode. When forking children, the process is copying the mmap to a a mallocated ring so we never share the same memory section between the master and the workers. When switching to wait mode, the shm is not used anymore as it is also copied to a mallocated structure. This allow to use the "show startup-logs" command over the master CLI, to get the logs of the latest startup or reload. This way the logs of the latest failed reload are also kept. This is only activated on the linux-glibc target for now.	2022-10-13 16:50:22 +02:00
William Lallemand	9e4ead3095	MINOR: ring: ring_cast_from_area() cast from an allocated area Cast an unified ring + storage area to a ring from area, without reinitializing the data buffer. Reinitialize the waiters and the lock. It helps retrieving a previously allocated ring, from an mmap for example.	2022-10-13 16:45:28 +02:00
Amaury Denoyelle	91b2305ad7	MINOR: quic: implement datagram cleanup for quic_receiver_buf Each time data is read on QUIC receiver socket, we try to reuse the first datagram of the currently used quic_receiver_buf instead of allocating a new one. This algorithm is suboptimal if there is several unused datagrams as only the first one is tested and its buffer removed from quic_receiver_buf. If QUIC traffic is quite substential, this can lead to an important number of quic_dgram occurences allocated from pool_head_quic_dgram and a lack of free space in allocated quic_receiver_buf buffers. To improve this, each time we want to reuse a datagram, we pop elements until a non-yet released datagram is found or the list is empty. All intermediary elements are freed and the last found datagram can be reused. This operation has been extracted in a dedicated function named quic_rxbuf_purge_dgrams(). This should improve memory consumption incured by quic_dgram instances under heavy QUIC traffic. Note that there is still room for improvement as if the first datagram is still in use, it may block several unused datagram after him. However this requires to support removal of datagrams out of order which is currently not possible. This should be backported up to 2.6.	2022-10-13 11:06:48 +02:00
Amaury Denoyelle	1cba8d60f3	CLEANUP: quic: improve naming for rxbuf/datagrams handling QUIC datagrams are read from a random thread. They are then redispatch to the connection thread according to the first packet DCID. These operations are implemented through a special buffer designed to avoid locking. Refactor this code with the following changes : * <rxbuf> type is renamed <quic_receiver_buf>. Its list element is also renamed to highligh its attach point to a receiver. * <quic_dgram> and <quic_receiver_buf> definition are moved to quic_sock-t.h. This helps to reduce the size of quic_conn-t.h. * <quic_dgram> list elements are renamed to highlight their attach point into a <quic_receiver_buf> and a <quic_dghdlr>. This should be backported up to 2.6.	2022-10-13 11:06:48 +02:00
Amaury Denoyelle	8c4d062d25	CLEANUP: quic: remove unused rxbufs member in receiver rxbuf is the structure used to store QUIC datagrams and redispatch them to the connection thread. Each receiver manages a list of rxbuf. This was stored both as an array and a mt_list. Currently, only mt_list is needed so removed <rxbufs> member from receiver structure. This should be backported up to 2.6.	2022-10-13 11:05:41 +02:00
Frédéric Lécaille	e1a49cfd4d	MINOR: quic: Split the secrets key allocation in two parts Implement quic_tls_secrets_keys_alloc()/quic_tls_secrets_keys_free() to allocate the memory for only one direction (RX or TX). Modify ha_quic_set_encryption_secrets() to call these functions for one of this direction (or both). So, for now on we can rely on the value of the secret keys to know if it was derived. Remove QUIC_FL_TLS_SECRETS_SET flag which is no more useful. Consequently, the secrets are dumped by the traces only if derived. Must be backported to 2.6.	2022-10-13 10:12:03 +02:00
Frédéric Lécaille	4aa7d8197a	BUG/MINOR: quic: Stalled 0RTT connections with big ClientHello TLS message This issue was reproduced with -Q picoquic client option to split a big ClientHello message into two Initial packets and haproxy as server without any knowledged of any previous ORTT session (restarted after a firt 0RTT session). The ORTT received packets were removed from their queue when the second Initial packet was parsed, and the QUIC handshake state never progressed and remained at Initial state. To avoid such situations, after having treated some Initial packets we always check if there are ORTT packets to parse and we never remove them from their queue. This will be done after the hanshake is completed or upon idle timeout expiration. Also add more traces to be able to analize the handshake progression. Tested with ngtcp2 and picoquic Must be backported to 2.6.	2022-10-13 10:12:03 +02:00
Frédéric Lécaille	9f9263ed13	MINOR: quic: Use a non-contiguous buffer for RX CRYPTO data Implement quic_get_ncbuf() to dynamically allocate a new ncbuf to be attached to any quic_cstream struct which needs such a buffer. Note that there is no quic_cstream for 0RTT encryption level. quic_free_ncbuf() is added to release the memory allocated for a non-contiguous buffer. Modify qc_handle_crypto_frm() to call this function and allocate an ncbuf for crypto data which are not received in order. The crypto data which are received in order are not buffered but provide to the TLS stack (calling qc_provide_cdata()). Modify qc_treat_rx_crypto_frms() which is called after having provided the in order received crypto data to the TLS stack to provide again the remaining crypto data which has been buffered, if possible (if they are in order). Each time buffered CRYPTO data were consumed, we try to release the memory allocated for the non-contiguous buffer (ncbuf). Also move rx.crypto.offset quic_enc_level struct member to rx.offset quic_cstream struct member. Must be backported to 2.6.	2022-10-13 10:12:03 +02:00
Frédéric Lécaille	a20c93e6e2	MINOR: quic: Extract CRYPTO frame parsing from qc_parse_pkt_frms() Implement qc_handle_crypto_frm() to parse a CRYPTO frame. Must be backported to 2.6.	2022-10-13 10:12:03 +02:00
Frédéric Lécaille	7e3f7c47e9	MINOR: quic: New quic_cstream object implementation Add new quic_cstream struct definition to implement the CRYPTO data stream. This is a simplication of the qcs object (QUIC streams) for the CRYPTO data without any information about the flow control. They are not attached to any tree, but to a QUIC encryption level, one by encryption level except for the early data encryption level (for 0RTT). A stream descriptor is also allocated for each CRYPTO data stream. Must be backported to 2.6	2022-10-13 10:12:03 +02:00
Willy Tarreau	d114f4a68f	MEDIUM: checks: spread the checks load over random threads The CPU usage pattern was found to be high (5%) on a machine with 48 threads and only 100 servers checked every second That was supposed to be only 100 connections per second, which should be very cheap. It was figured that due to the check tasks unbinding from any thread when going back to sleep, they're queued into the shared queue. Not only this requires to manipulate the global queue lock, but in addition it means that all threads have to check the global queue before going to sleep (hence take a lock again) to figure how long to sleep, and that they would all sleep only for the shortest amount of time to the next check, one would pick it and all other ones would go down to sleep waiting for the next check. That's perfectly visible in time-to-first-byte measurements. A quick test consisting in retrieving the stats page in CSV over a 48-thread process checking 200 servers every 2 seconds shows the following tail: percentile ttfb(ms) 99.98 2.43 99.985 5.72 99.99 32.96 99.995 82.176 99.996 82.944 99.9965 83.328 99.997 83.84 99.9975 84.288 99.998 85.12 99.9985 86.592 99.999 88 99.9995 89.728 99.9999 100.352 One solution could consist in forcefully binding checks to threads at boot time, but that's annoying, will cause trouble for dynamic servers and may cause some skew in the load depending on some server patterns. Instead here we take a different approach. A check remains bound to its thread for as long as possible, but upon every wakeup, the thread's load is compared with another random thread's load. If it's found that that other thread's load is less than half of the current one's, the task is bounced to that thread. In order to prevent that new thread from doing the same, we set a flag "CHK_ST_SLEEPING" that indicates that it just woke up and we're bouncing the task only on this condition. Tests have shown that the initial load was very unfair before, with a few checks threads having a load of 15-20 and the vast majority having zero. With this modification, after two "inter" delays, the load is either zero or one everywhere when checks start. The same test shows a CPU usage that significantly drops, between 0.5 and 1%. The same latency tail measurement is much better, roughly 10 times smaller: percentile ttfb(ms) 99.98 1.647 99.985 1.773 99.99 4.912 99.995 8.76 99.996 8.88 99.9965 8.944 99.997 9.016 99.9975 9.104 99.998 9.224 99.9985 9.416 99.999 9.8 99.9995 10.04 99.9999 10.432 In fact one difference here is that many threads work while in the past they were waking up and going down to sleep after having perturbated the shared lock. Thus it is anticipated that this will scale way smoother than before. Under strace it's clearly visible that all threads are sleeping for the time it takes to relaunch a check, there's no more thundering herd wakeups. However it is also possible that in some rare cases such as very short check intervals smaller than a scheduler's timeslice (such as 4ms), some users might have benefited from the work being concentrated on less threads and would instead observe a small increase of apparent CPU usage due to more total threads waking up even if that's for less work each and less total work. That's visible with 200 servers at 4ms where show activity shows that a few threads were overloaded and others doing nothing. It's not a problem, though as in practice checks are not supposed to eat much CPU and to wake up fast enough to represent a significant load anyway, and the main issue they could have been causing (aside the global lock) is an increase last-percentile latency.	2022-10-12 21:49:30 +02:00
Willy Tarreau	a840b4a39b	MINOR: checks: use the lighter PRNG for spread checks There's no point using ha_random32() which is heavy and uses shared variables to calculate a random timer when we have statistical_prng() which does the same and was made exactly for this.	2022-10-12 21:49:30 +02:00
Willy Tarreau	99521abd59	BUG/MINOR: server: make sure "show servers state" hides private bits In the past we've seen "show servers state" dump some internal bits for the check states, that were causing regtests to fail. The relevant bits have been added to the doc to fix the public API and make sure they do not change by accident, but the output doesn't take care of masking the undesired ones, causing regtests (and possibly user programs) to fail when new bits are added. Let's add the mask for the only documented ones (0x0F for check and 0x1F for agent respectively). This could be backported wherever the server state is present, though there's a tiny risk that some undocumented bits might have already leaked to some user scripts, so it might be wise to wait a bit before doing that or even not to backport too far.	2022-10-12 21:45:39 +02:00
Christopher Faulet	104985610d	BUG/MEDIUM: mux-h1: Handle abort with an incomplete message during parsing In h1_process_demux(), aborts for incomplete messages were not properly handled. It was not an issue because the abort was detected later in h1_process(). But it will be an issue to perform the aborts refoctoring. First, when a read0 was detected, the SE_FL_EOI flag was set for messages in DONE or TUNNEL state or for messages without known length (so responses in close mode). The last statement is not accurate. The message must also be in DATA state. Otherwise, SE_FL_EOI flag may be set on incomplete message. Then, an error was reported, via SE_FL_ERROR flag, only when an incomplete message was detected on the payload parsing. It must also be reported if headers are incomplete. Here again, the error is detected later for now. But it could be an issue later. There is no reason to backport this patch.	2022-10-12 17:10:44 +02:00
Christopher Faulet	9009c974c1	BUG/MEDIUM: mux-h1: Add connection error handling when reading/sending on a pipe There is no error handling when we read or write on a pipe. There error is caught later, in the mux I/O handler. But there is no reason to not do so here. There is no reason to backport it because no issue was reported for now because of this "bug". In all cases, it must be evaluated first.	2022-10-12 17:10:44 +02:00
Christopher Faulet	3965aa7494	REORG: mux-fcgi: Extract flags and enums into mux_fcgi-t.h The same was performed for the H2 and H1 multiplexers. FCGI connection and stream flags are moved in a dedicated header file. It will be mainly used to be able to decode mux-fcgi flags from the flags utility. In this patch, we move the flags and enums to mux_fcgi-t.h, as well as the two state decoding inline functions.	2022-10-12 17:10:37 +02:00
Amaury Denoyelle	3e0648837c	BUG/MINOR: stick-table: fix build with DEBUG_THREAD Compilation is broken with DEBUG_THREAD since the following patch `76642223f0` MEDIUM: stick-table: switch the table lock to rwlock Fix this by updating a legacy HA_SPIN_INIT() to HA_RWLOCK_INIT(). No backport needed unless the mentionned patch is backported.	2022-10-12 16:54:59 +02:00
Willy Tarreau	cbdb528a76	MEDIUM: stick-table: requeue the wakeup task out of the write lock We don't need to call stktable_requeue_exp() with the table's lock held anymore, so let's move it out. It should slightly reduce the contention on the write lock, though it is now already quite low.	2022-10-12 14:19:05 +02:00
Willy Tarreau	dbae89e09c	MEDIUM: stick-table: always use atomic ops to requeue the table's task We're generalizing the change performed in previous commit "MEDIUM: stick-table: requeue the expiration task out of the exclusive lock" to stktable_requeue_exp() so that it can also be used by callers of __stktable_store(). At the moment there's still no visible change since it's still called under the write lock. However, the previous code in stitable_touch_with_exp() was updated to use this function.	2022-10-12 14:19:05 +02:00
Willy Tarreau	eb23e3e243	MINOR: stick-table: split stktable_store() between key and requeue __staktable_store() performs two distinct things, one is to insert a key and the other one is to requeue the task's expiration date. Since the latter might be done without a lock, let's first split the function in two halves. For now this has no impact.	2022-10-12 14:19:05 +02:00
Willy Tarreau	e3f5ae895a	MEDIUM: stick-table: requeue the expiration task out of the exclusive lock With 48 threads, a heavily loaded table with plenty of trackers and rules and a short expiration timer of 10ms saturates the CPU at 232k rps. By carefully using atomic ops we can make sure that t->exp_next and t->task->expire converge to the earliest next expiration date and that all of this can be performed under atomic ops without any lock. That's what this patch is doing in stktable_touch_with_exp(). This is sufficient to double the performance and reach 470k rps. It's worth noting that __stktable_store() uses a mix of eb32_insert() and task_queue, and that the second part of it could possibly benefit from this, even though sometimes it's called under a lock that was already held.	2022-10-12 14:19:05 +02:00
Willy Tarreau	e62885237c	MEDIUM: stick-table: make stktable_set_entry() look up under a read lock On a 24-core machine having some "stick-store response" rules, a lot of time is spent in the write lock in stktable_set_entry(). Let's apply the same mechanism as for the stktable_get_entry() consisting in looking up the value under the read lock and upgrading it to a write lock only to perform modifications. Here we even have the luxury of upgrading the lock since there are no alloc/free in the path. All this increases the performance by 40% (from 363k to 510k rps).	2022-10-12 14:19:05 +02:00
Willy Tarreau	996f1a5124	MEDIUM: stick-table: do not take a lock to update t->current anymore. We don't need to be protected by the table's lock when touching t->current if we do it using atomics, and that's great because it allows us to have a cleaner stksess_new() that doesn't require a lock either, and to avoid manipulating pools under a lock. That's another 1% performance gain from 2.07 to 2.10M req/s under 48 threads.	2022-10-12 14:19:05 +02:00
Willy Tarreau	47f229702e	MEDIUM: stick-table: make stktable_get_entry() look up under a read lock On a 24-core machine doing lots of track-sc, it was found that the lock in stktable_get_entry() was responsible for 25% of the CPU alone. It's sad because most of its job is to protect the table during the lookup. Here we're taking a slightly different approach: the lock is first taken for reads during the lookup, and only in case of failure we switch it for a write lock. We don't even perform an upgrade here since an allocation is needed between the two, it would be wasted to do it under the lock, and is generally not a good idea, so better release the read lock and try again. Here the performance under 48 threads with 3 trackers on the same table jumped from 455k to 2.07M, or 4.55x! Note that the same approach should be possible for stktable_set_entry().	2022-10-12 14:19:05 +02:00
Willy Tarreau	a7d6a1396e	MEDIUM: stick-table: switch to rdlock in stktable_lookup() and lookup_key() These functions do not modify anything in the the table except the refcount on success. Let's just lock the table for shared accesses and make use of atomic ops to update the refcount. This brings a nice gain from 425k to 455k under 48 threads (7%), but some contention remains on the exclusive locks in other parts. Note that the refcount continues to be updated under the lock because it's not yet certain whether there are races between it and some of the exclusive lock on the table. The difference is marginal and we prefer to stay on the safe side for now.	2022-10-12 14:19:05 +02:00
Willy Tarreau	175aa06232	MEDIUM: stick-table: free newly allocated stkess if it couldn't be inserted In __stktable_get_entry() now we're planning for the possibility that the call to __stktable_store() doesn't add the newly allocated entry and instead finds a previously inserted one. At the moment this doesn't exist because the lookup + insert passes are made under the same lock. But it will soon change.	2022-10-12 14:19:05 +02:00
Willy Tarreau	d2d3fd9b5e	MEDIUM: stick-table: return inserted entry in __stktable_store() This function is used to create an entry in the table. But it doesn't consider the possibility that the entry already exists, because right now it's only called in situations where it was verified under a lock that it doesn't exist. Since we'll soon need to break that assumption we need it to verify that the requested entry was added and to return a pointer to the one in the tree so that the caller can detect any possible conflict. At the moment this is not used.	2022-10-12 14:19:05 +02:00
Willy Tarreau	8d3c3336f9	MEDIUM: stick-table: make stksess_kill_if_expired() avoid the exclusive lock stream_store_counters() calls stksess_kill_if_expired() for each active counter. And this one takes an exclusive lock on the table before checking if it has any work to do (hint: it almost never has since it only wants to delete expired entries). However a lock is still neeed for now to protect the ref_cnt, but we can do it atomically under the read lock. Let's change the mechanism. Now what we do is to check out of the lock if the entry is expired. If it is, we take the write lock, expire it, and decrement the refcount. Otherwise we just decrement the refcount under a read lock. With this change alone, the config based on 3 trackers without the previous patches saw a 2.6x improvement, but here it doesn't yet change anything because some heavy contention remains on the lookup part.	2022-10-12 14:19:05 +02:00
Willy Tarreau	a7536ef9e1	MEDIUM: stick-table: only take the lock when needed in stktable_touch_with_exp() As previously mentioned, this function currently holds an exclusive lock on the table during all the time it take to check if the entry needs to be updated and synchronized with peers. The reality is that many setups do not use peers and that on highly loaded setups, the same entries are hammered all the time so the key's expiration doesn't change between a number of consecutive accesses. With this patch we take a different approach. The function starts without taking the lock, and will take it only if needed, keeping track of it. This way we can avoid it most of the time, or even entirely. Finally if the decrefcnt argument requires that the refcount is decremented, we either do it using a non-atomic op if the table was locked (since no other entry may touch it) or via an atomic under the read lock only. With this change alone, a 48-thread test with 3 trackers increased from 193k req/s to 425k req/s, which is a 2.2x factor.	2022-10-12 14:19:05 +02:00
Willy Tarreau	9f5cb435b6	MINOR: stick-table: move the write lock inside stktable_touch_with_exp() Taking the write lock prior to entering that function is a problem because this function is full of conditions that most of the time can lead to eliminating the lock. This commit first moves the write lock inside the function and passes the extra argument required to implement stktable_touch_remote() and stktable_touch_local(). It also renames the function to remove the underscores since there's no other variant and it's exported under this name (probably an old rename that was not propagated). The code was stressed under 48 threads using 3 trackers on the same table. It already shows a tiny 3% improvement from 187k to 193k rps.	2022-10-12 14:19:05 +02:00
Willy Tarreau	4be073b99b	MINOR: stick-table: do not take an exclusive lock when downing ref_cnt At plenty of places we decrement ts->ref_cnt under the write lock because it's held. We don't technically need it to be done that way if there's contention and an atomic could suffice. However until all places are turned to atomic, we at least need to do that under a read lock for now, so that we don't mix atomic and non-atomic uses. Regardless it already brings ~1.5% req rate improvement with 3 trackers on the same table under 48 threads at 184k->187k rps.	2022-10-12 14:19:05 +02:00
Willy Tarreau	76642223f0	MEDIUM: stick-table: switch the table lock to rwlock Right now a spinlock is used, but most accesses are for reads, so let's switch the lock to an rwlock and switch all accesses to exclusive locks for now. There should be no visible difference at this point.	2022-10-12 14:19:05 +02:00
Willy Tarreau	f6a42c3a37	MINOR: freq_ctr: use the thread's local time whenever possible Right now when dealing with freq_ctr updates, we're using the process- wide monotinic time, and accessing it is expensive since every thread needs to update it, so this adds some contention. However we don't need it all the time, the thread's local time is most of the time strictly equal to the global time, and may be off by one millisecond when the global time is switched to the next one by another thread, and in this case we don't want to use the local time because it would risk to cause a rotation of the counter. But that's precisely the condition we're already relying on for the slow path! What this patch does is to add a check for the period against the local time prior to anything else, and immediately return after updating the counter if still within the period, otherwise fall back to the existing code. Given that the function starts to inflate a bit, it was split between s very short inline part that does the hot path, and the slower fallback that's in a cold function. It was measured that on a 24-CPU machine it was called ~0.003% of the time. The resulting improvement sits between 2 and 3% at 500k req/s tracking an http_req_rate counter.	2022-10-12 14:19:05 +02:00
Willy Tarreau	bc7c207f74	BUG/MAJOR: stick-tables: do not try to index a server name for applets Since commit `03cdf55e6` ("MINOR: stream: Stickiness server lookup by name.") in 2.0-dev6, server names may be used instead of their IDs, in order to perform stickiness. However the commit above may end up trying to insert an empty server name in the dictionary when the server is an applet instead, resulting in an immediate segfault. This is typically what happens when a "stick-store" rule is present in a backend featuring a "stats" directive. As there doesn't seem to be an easy way around it, it seems to imply that "stick-store" is not much used anymore. The solution here is to only try to insert non-null keys into the dictionary. The patch moves the check of the key type before the first lock so that the test on the key can be performed under the lock instead of locking twice (the patch is more readable with diff -b). Note that before 2.4, there's no <key> variable there as it was introduced by commit `92149f9a8` ("MEDIUM: stick-tables: Add srvkey option to stick-table"), but the __objt_server(s->target)->id still needs to be tested. This needs to be backported as far as 2.0.	2022-10-12 14:19:05 +02:00
Aurelien DARRAGON	d56bebee7b	MINOR: hlua: removing ambiguous lua_pushvalue with 0 index In `cd341d531`, I added a FIXME comment because I noticed a lua_pushvalue with 0 index, whereas lua doc states that 0 is never an acceptable index. After reviewing and testing the hlua_applet_http_send_response() code, it turns out that this pushvalue is not even needed. So it's safer to remove it as it could lead to undefined behavior (since it is not supported by Lua API) and it grows lua stack by 1 for no reason. No backport needed.	2022-10-12 09:22:05 +02:00
Aurelien DARRAGON	d83d045cda	MINOR: hlua: some luaL_checktype() calls were not guarded with MAY_LJMP In hlua code, we mark every function that may longjump using MAY_LJMP macro so it's easier to identify them by reading the code. However, some luaL_checktypes() were performed without the MAY_LJMP. According to lua doc: Functions called luaL_check* always raise an error if the check is not satisfied. -> Adding the missing MAY_LJMP for those luaLchecktypes() calls. No backport needed.	2022-10-12 09:22:05 +02:00
Amaury Denoyelle	487d04f6d7	BUG/MINOR: quic: set IP_PKTINFO socket option for QUIC receivers only Move code which activates IP_PKTINFO socket option (or affiliated options) from sock_inet_bind_receiver() to quic_bind_listener() function. This change is useful for two reasons : * first, and the most important one : this activates IP_PKTINFO only for QUIC receivers. The previous version impacted all datagram receivers, used for example by log-forwarder. This should reduce memory usage for these datagram sockets which do not need this option. * second, USE_QUIC preprocessor statements are removed from src/sock_inet.c which clean up the code. IP_PKTINFO was introduced recently by the following patch : `97ecc7a8ea` (quic-dev/qns) MEDIUM: quic: retrieve frontend destination address For the moment, this does not impact any stable release. However, as previous patch is scheduled for 2.6 backporting, the current change must also be backported to the same versions.	2022-10-11 16:46:04 +02:00
Willy Tarreau	cab054bbf9	CLEANUP: quic/receiver: remove the now unused tx_qring list The tx_qrings[] and tx_qring_list in the receiver are not used anymore since commit `f2476053f` ("MINOR: quic: replace custom buf on Tx by default struct buffer"), the only place where they're referenced was in quic_alloc_tx_rings_listener(), which by the way implies that these were not even freed on exit. Let's just remove them. This should be backported to 2.6 since the commit above also was.	2022-10-11 08:40:38 +02:00
Tim Duesterhus	a6fc616c1a	CLEANUP: Reapply strcmp.cocci This reapplies strcmp.cocci across the whole src/ tree.	2022-10-10 15:49:09 +02:00
Tim Duesterhus	a029d781e2	CLEANUP: Reapply ist.cocci (2) This reapplies ist.cocci across the whole src/ tree.	2022-10-10 15:49:09 +02:00
Amaury Denoyelle	97ecc7a8ea	MEDIUM: quic: retrieve frontend destination address Retrieve the frontend destination address for a QUIC connection. This address is retrieve from the first received datagram and then stored in the associated quic-conn. This feature relies on IP_PKTINFO or affiliated flags support on the socket. This flag is set for each QUIC listeners in sock_inet_bind_receiver(). To retrieve the destination address, recvfrom() has been replaced by recvmsg() syscall. This operation and parsing of msghdr structure has been extracted in a wrapper quic_recv(). This change is useful to finalize the implementation of 'dst' sample fetch. As such, quic_sock_get_dst() has been edited to return local address from the quic-conn. As a best effort, if local address is not available due to kernel non-support of IP_PKTINFO, address of the listener is returned instead. This should be backported up to 2.6.	2022-10-10 11:48:27 +02:00
Amaury Denoyelle	90121b3321	CLEANUP: quic: fix indentation Fix some indentation in qc_lstnr_pkt_rcv(). This should be backported up to 2.6.	2022-10-05 11:08:32 +02:00
Amaury Denoyelle	036cc5d880	MINOR: mux-quic: check quic-conn return code on Tx Inspect return code of qc_send_mux(). If quic-conn layer reports an error, this will interrupt the current emission process. This should be backported up to 2.6.	2022-10-05 11:08:32 +02:00
Amaury Denoyelle	2ed840015f	MINOR: quic: limit usage of ssl_sock_ctx in favor of quic_conn Continue on the cleanup of QUIC stack and components. quic_conn uses internally a ssl_sock_ctx to handle mandatory TLS QUIC integration. However, this is merely as a convenience, and it is not equivalent to stackable ssl xprt layer in the context of HTTP1 or 2. To better emphasize this, ssl_sock_ctx usage in quic_conn has been removed wherever it is not necessary : namely in functions not related to TLS. quic_conn struct now contains its own wait_event for tasklet quic_conn_io_cb(). This should be backported up to 2.6.	2022-10-05 11:08:32 +02:00
Aurelien DARRAGON	afb7dafb44	BUG/MINOR: hlua: hlua_channel_insert_data() behavior conflicts with documentation Channel.insert(channel, string, [,offset]): When no offset is provided, hlua_channel_insert_data() inserts string at the end of incoming data. This behavior conflicts with the documentation that explicitly says that the default behavior is to insert the string in front of incoming data. This patch fixes hlua_channel_insert_data() behavior so that it fully complies with the documentation. Thanks to Smackd0wn for noticing it. This could be backported to 2.6 and 2.5	2022-10-05 11:03:56 +02:00
Willy Tarreau	2e2b79d157	BUILD: http_fetch: silence an uninitiialized warning with gcc-4/5/6 at -Os Gcc 4.x, 5.x and 6.x report this when compiling http_fetch.c: src/http_fetch.c: In function 'smp_fetch_meth': src/http_fetch.c:357:6: warning: 'htx' may be used uninitialized in this function [-Wmaybe-uninitialized] sl = http_get_stline(htx); That's quite weird since there's no such code path, but presetting the htx variable to NULL during declaration is enough to shut it up. This may be backported to any version that has `dbbdb25f1` ("BUG/MINOR: http-fetch: Use integer value when possible in "method" sample fetch") as it's the one that triggered this warning (hence at least 2.0).	2022-10-04 09:18:34 +02:00
Christopher Faulet	eefcd8a97d	BUG/MINOR: http-fetch: Update method after a prefetch in smp_fetch_meth() In smp_fetch_meth(), smp_prefetch_htx() function may be called to parse the HTX message and update the HTTP transaction accordingly. In this case, in smp_fetch_metch() and on success, we must update "meth" variable. Otherwise, the variable is still equal to HTTP_METH_OTHER and the string version is always used instead of the enum for known methods. This patch must be backported as far as 2.0.	2022-10-04 09:16:36 +02:00
Willy Tarreau	c06557c23b	MINOR: init: do not try to shrink existing RLIMIT_NOFIlE As seen in issue #1866, some environments will not allow to change the current FD limit, and actually we don't need to do it, we only do it as a byproduct of adjusting the limit to the one that fits. Here we're replacing calls to setrlimit() with calls to raise_rlim_nofile(), which will avoid making the setrlimit() syscall in case the desired value is lower than the current process' one. This depends on previous commit "MINOR: fd: add a new function to only raise RLIMIT_NOFILE" and may need to be backported to 2.6, possibly earlier, depending on users' experience in such environments.	2022-10-04 08:38:47 +02:00
Willy Tarreau	922a907926	MINOR: fd: add a new function to only raise RLIMIT_NOFILE In issue #1866 an issue was reported under docker, by which a user cannot lower the number of FD needed. It looks like a restriction imposed in this environment, but it results in an error while it ought not have to in the case of shrinking. This patch adds a new function raise_rlim_nofile() that takes the desired new setting, compares it to the current one, and only calls setrlimit() if one of the values in the new setting is larger than the older one. As such it will continue to emit warnings and errors in case of failure to raise the limit but will never shrink it. This patch is only preliminary to another one, but will have to be backported where relevant (likely only 2.6).	2022-10-04 08:38:47 +02:00
Willy Tarreau	55d2e8577e	BUILD: h1: silence an initiialized warning with gcc-4.7 and -Os Building h1.c with gcc-4.7 -Os produces the following warning: src/h1.c: In function 'h1_headers_to_hdr_list': src/h1.c:1101:36: warning: 'ptr' may be used uninitialized in this function [-Wmaybe-uninitialized] In fact ptr may be taken from sl.rq.u.ptr which is only initialized after passing through the relevant states, but gcc doesn't know which states are visited. Adding an ALREADY_CHECKED() statement there is sufficient to shut it up and doesn't affect the emitted code. This may be backported to stable versions to make sure that builds on older distros and systems is clean.	2022-10-04 08:02:03 +02:00
Olivier Houchard	14f6268883	BUG/MEDIUM: lua: handle stick table implicit arguments right. In hlua_lua2arg_check(), we allow for the first argument to not be provided, if it has a type we know, this is true for frontend, backend, and stick table. However, the stick table code was changed. It used to be deduced from the proxy, but it is now directly provided in struct args. So setting the proxy there no longer work, and we have to explicitely set the stick table. Not doing so will lead the code do use the proxy pointer as a stick table pointer, which will likely cause crashes. This should be backported up to 2.0.	2022-10-03 19:08:10 +02:00
Olivier Houchard	ca43161a8d	BUG/MEDIUM: lua: Don't crash in hlua_lua2arg_check on failure In hlua_lua2arg_check(), on failure, before calling free_argp(), make sure to always mark the failed argument as ARGT_STOP. We only want to free argument prior to that point, because we did not allocate the strings after this one, and so we don't want to free them. This should be backported up to 2.2.	2022-10-03 19:08:10 +02:00
Amaury Denoyelle	d7755375a5	BUG/MINOR: mux-quic: ignore STOP_SENDING for locally closed stream It is possible to receive a STOP_SENDING frame for a locally closed stream. This was not properly managed as this would result in a BUG_ON() crash from qcs_idle_open() call under qcc_recv_stop_sending(). Now, STOP_SENDING frames are ignored when received on streams already locally closed. This has two consequences depending on the reason of closure : * if a RESET_STREAM was already emitted and closed the stream, this patch prevents to emit a new RESET_STREAM. This behavior is thus better. * if stream was closed due to all data transmitted, no RESET_STREAM will be built. This is contrary to the RFC 9000 which advice to transmit it, even on "Data Sent" state. However, this is not mandatory so the new behavior is acceptable, even if it could be improved. This crash has been detected on haproxy.org. This can be artifically reproduced by adding the following snippet at the end of qc_send_mux() when doing a request with a small payload response : qcc_recv_stop_sending(qc->qcc, 0, 0); This must be backported up to 2.6.	2022-10-03 17:20:31 +02:00
Amaury Denoyelle	92fa63f735	CLEANUP: quic: create a dedicated quic_conn module xprt_quic module was too large and did not reflect the true architecture by contrast to the other protocols in haproxy. Extract code related to XPRT layer and keep it under xprt_quic module. This code should only contains a simple API to communicate between QUIC lower layer and connection/MUX. The vast majority of the code has been moved into a new module named quic_conn. This module is responsible to the implementation of QUIC lower layer. Conceptually, it overlaps with TCP kernel implementation when comparing QUIC and HTTP1/2 stacks of haproxy. This should be backported up to 2.6.	2022-10-03 16:25:17 +02:00
Amaury Denoyelle	ac9bf016bf	CLEANUP: quic: remove unused function prototype Removed hexdump unusued prototype from quic_tls.c. This should be backported up to 2.6.	2022-10-03 16:25:17 +02:00
Amaury Denoyelle	5c25dc5bfd	CLEANUP: quic: fix headers Clean up quic sources by adjusting headers list included depending on the actual dependency of each source file. On some occasion, xprt_quic.h was removed from included list. This is useful to help reducing the dependency on this single file and cleaning up QUIC haproxy architecture. This should be backported up to 2.6.	2022-10-03 16:25:17 +02:00
Amaury Denoyelle	f3c40f83fb	BUG/MINOR: quic: adjust quic_tls prototypes Two prototypes in quic_tls module were not identical to the actual function definition. * quic_tls_decrypt2() : the second argument const attribute is not present, to be able to use it with EVP_CIPHER_CTX_ctlr(). As a consequence of this change, token field of quic_rx_packet is now declared as non-const. * quic_tls_generate_retry_integrity_tag() : the second argument type differ between the two. Adjust this by fixing it to as unsigned char to match EVP_EncryptUpdate() SSL function. This situation did not seem to have any visible effect. However, this is clearly an undefined behavior and should be treated as a bug. This should be backported up to 2.6.	2022-10-03 16:25:17 +02:00
Amaury Denoyelle	a19bb6f0b2	CLEANUP: quic: remove global var definition in quic_tls header Some variables related to QUIC TLS were defined in a header file : their definitions are now moved properly in the implementation file, with only declarations in the header. This should be backported up to 2.6.	2022-10-03 16:25:17 +02:00
Amaury Denoyelle	d6922d5471	CLEANUP: mux-quic: remove usage of non-standard ull type ull is a typedef to unsigned long long. It is only defined in xprt_quic-t.h. Its usage should be limited over time to reduce xprt_quic dependency over the whole code. It can be replaced by ullong typedef from compat.h. For the moment, ull references have been replaced in qmux_trace module. They were only used for printf format and has been replaced by the true variable type. This change is useful to reduce dependencies on xprt_quic in other files. This should be backported up to 2.6.	2022-10-03 16:24:44 +02:00
Fatih Acar	0d6fb7a3eb	BUG/MINOR: checks: update pgsql regex on auth packet This patch adds support to the following authentication methods: - AUTH_REQ_GSS (7) - AUTH_REQ_SSPI (9) - AUTH_REQ_SASL (10) Note that since AUTH_REQ_SASL allows multiple authentication mechanisms such as SCRAM-SHA-256 or SCRAM-SHA-256-PLUS, the auth payload length may vary since the method is sent in plaintext. In order to allow this, the regex now matches any payload length. This partially fixes Github issue #1508 since user authentication is still broken but should restore pre-2.2 behavior. This should be backported up to 2.2. Signed-off-by: Fatih Acar <facar@scaleway.com>	2022-10-03 15:31:22 +02:00
Willy Tarreau	406efb96d1	BUG/MINOR: backend: only enforce turn-around state when not redispatching In github issue #1878, Bart Butler reported observing turn-around states (1 second pause) after connection retries going to different servers, while this ought not happen. In fact it does happen because back_handle_st_cer() enforces the TAR state for any algo that's not round-robin. This means that even leastconn has it, as well as hashes after the number of servers changed. Prior to doing that, the call to stream_choose_redispatch() has already had a chance to perform the correct choice and to check the algo and the number of retries left. So instead we should just let that function deal with the algo when needed (and focus on deterministic ones), and let the former just obey. Bart confirmed that the fixed version works as expected (no more delays during retries). This may be backported to older releases, though it doesn't seem very important. At least Bart would like to have it in 2.4 so let's go there for now after it has cooked a few weeks in 2.6.	2022-10-03 15:04:55 +02:00
Thierry Fournier	3d1c334d44	BUG/MINOR: config: insufficient syntax check of the global "maxconn" value The maxconn value is decoded using atol(), so values like "3k" are rightly converter as interger 3, while the user wants 3000. This patch fixes this behavior by reporting a parsing error. This patch could be backported on all maintained version, but it could break some configuration. The bug is really minor, I recommend to not backport, or backport a patch which only throws a warning in place of a fatal error.	2022-10-03 14:30:08 +02:00
Willy Tarreau	8522348482	BUG/MAJOR: conn-idle: fix hash indexing issues on idle conns Idle connections do not work on 32-bit machines due to an alignment issue causing the connection nodes to be indexed with their lower 32-bits set to zero and the higher 32 ones containing the 32 lower bitss of the hash. The cause is the use of ebmb_node with an aligned data, as on this platform ebmb_node is only 32-bit aligned, leaving a hole before the following hash which is a uint64_t: $ pahole -C conn_hash_node ./haproxy struct conn_hash_node { struct ebmb_node node; /* 0 20 / / XXX 4 bytes hole, try to pack / int64_t hash; / 24 8 / struct connection conn; /* 32 4 / / size: 40, cachelines: 1, members: 3 / / sum members: 32, holes: 1, sum holes: 4 / / padding: 4 / / last cacheline: 40 bytes */ }; Instead, eb64 nodes should be used when it comes to simply storing a 64-bit key, and that is what this patch does. For backports, a variant consisting in simply marking the "hash" member with a "packed" attribute on the struct also does the job (tested), and might be preferable if the fix is difficult to adapt. Only 2.6 and 2.5 are affected by this.	2022-10-03 12:06:36 +02:00
Willy Tarreau	94ab139266	BUG/MEDIUM: config: count line arguments without dereferencing the output Previous commit `8a6767d26` ("BUG/MINOR: config: don't count trailing spaces as empty arg (v2)") was still not enough. As reported by ClusterFuzz in issue 52049 (https://bugs.chromium.org/p/oss-fuzz/issues/detail?id=52049), there remains a case where for the sake of reporting the correct argument count, the function may produce virtual args that span beyond the end of the output buffer if that one is too short. That's what's happening with a config file of one empty line followed by a large number of args. This means that what args[] points to cannot be relied on and that a different approach is needed. Since no output is produced for spaces and comments, we know that args[arg] continues to point to out+outpos as long as only comments or spaces are found, which is what we're interested in. As such it's safe to check the last arg's pointer against the one before the trailing zero was emitted, in order to decide to count one final arg. No backport is needed, unless the commit above is backported.	2022-10-03 09:24:26 +02:00
Erwan Le Goas	8a6767d266	BUG/MINOR: config: don't count trailing spaces as empty arg (v2) In parse_line(), spaces increment the arg count and it is incremented again on '#' or end of line, resulting in an extra empty arg at the end of arg's list. The visible effect is that the reported arg count is in excess of 1. It doesn't seem to affect regular function but specialized ones like anonymisation depends on this count. This is the second attempt for this problem, here the explanation : When called for the first line, no <out> was allocated yet so it's NULL, letting the caller realloc a larger line if needed. However the words are parsed and their respective args[arg] are filled with out+position, which means that while the first arg is NULL, the other ones are no and fail the test that was meant to avoid dereferencing a NULL. Let's simply check <out> instead of <args> since the latter is always derived from the former and cannot be NULL without the former also being. This may need to be backported to stable versions.	2022-09-30 15:21:20 +02:00
Aurelien DARRAGON	cd341d5314	MINOR: hlua: ambiguous lua_pushvalue with 0 index In function hlua_applet_http_send_response(), a pushvalue is performed with index '0'. But according to lua doc (https://www.lua.org/manual/5.3/manual.html#4.3): "Note that 0 is never an acceptable index". Adding a FIXME comment near to the pushvalue operation so that this can get some chance to be reviewed later. No backport needed.	2022-09-30 15:21:20 +02:00
Aurelien DARRAGON	4d7aefeee1	BUG/MINOR: hlua: prevent crash when loading numerous arguments using lua-load(per-thread) When providing multiple optional arguments with lua-load or lua-load-per-thread directives, arguments where pushed 1 by 1 to the stack using lua_pushstring() without checking if the stack could handle it. This could easily lead to program crash when providing too much arguments. I can easily reproduce the crash starting from ~50 arguments. Calling lua_checkstack() before pushing to the stack fixes the crash: According to lua.org, lua_checkstack() does some housekeeping and allow the stack to be expanded as long as some memory is available and the hard limit isn't reached. When no memory is available to expand the stack or the limit is reached, lua_checkstacks returns an error: in this case we force hlua_load_state() to return a meaningfull error instead of crashing. In practice though, cfgparse complains about too many words way before such event may occur on a normal system. TLDR: the ~50 arguments limitation is not an issue anymore. No backport needed, except if 'MINOR: hlua: Allow argument on lua-lod(-per-thread) directives' (`ae6b568`) is backported.	2022-09-30 15:21:20 +02:00
Aurelien DARRAGON	bcbcf98e0c	BUG/MINOR: hlua: _hlua_http_msg_delete incorrect behavior when offset is used Calling the function with an offset when "offset + len" was superior or equal to the targeted blk length caused 'v' value to be improperly set. And because 'v' is directly provided to htx_replace_blk_value(), blk consistency was compromised. (It seems that blk was overrunning in htx_replace_blk_value() due to this and header data was overwritten in this case). This patch adds the missing checks to make the function behave as expected when offset is set and offset+len is greater or equals to the targeted blk length. Some comments were added to the function as well. It may be backported to 2.6 and 2.5	2022-09-29 12:03:04 +02:00
Christopher Faulet	015bbc298f	MINOR: tools: Impprove hash_ipanon to not hash FD-based addresses "stdout" and "stderr" are not hashed. In the same spirit, "fd@" and "sockpair@" prefixes are not hashed too. There is no reason to hash such address and it may be useful to diagnose bugs. No backport needed, except if anonymization mechanism is backported.	2022-09-29 11:53:08 +02:00
Christopher Faulet	7e50e4b9cc	MINOR: tools: Impprove hash_ipanon to support dgram sockets and port offsets Add PA_O_DGRAM and PA_O_PORT_OFS options when str2sa_range() is called. This way dgram sockets and addresses with port offsets are supported. No backport needed, except if anonymization mechanism is backported.	2022-09-29 11:46:35 +02:00
Erwan Le Goas	d78693178c	MINOR: cli: correct commentary and replace 'set global-key' name Correct a commentary in in include/haproxy/global-t.h and include/haproxy/tools.h Replace the CLI command 'set global-key <key>' by 'set anon global-key <key>' in order to find it easily when you don't remember it, the recommandation can guide you when you just tap 'set anon'. No backport needed, except if anonymization mechanism is backported.	2022-09-29 10:53:15 +02:00
Erwan Le Goas	f30c5d7666	MINOR: config: Add option line when the configuration file is dumped Add an option to dump the number lines of the configuration file when it's dumped. Other options can be easily added. Options are separated by ',' when tapping the command line: './haproxy -dC[key],line -f [file]' No backport needed, except if anonymization mechanism is backported.	2022-09-29 10:53:15 +02:00
Erwan Le Goas	059d05f702	MINOR: config: Add other keywords when dump the anonymized configuration file Add keywords recognized during the dump of the configuration file, these keywords are followed by sensitive information. Remove the condition 'localhost' for the second argument of keyword 'server', consider as not essential and can disturb when comparing it in cli section (there is no exception 'localhost'). No backport needed, except if anonymization mechanism is backported.	2022-09-29 10:53:15 +02:00
Erwan Le Goas	be5ed92d0a	MINOR: config: correct errors about argument number in condition in cfgparse.c Put the right number in condition that takes the wrong number of arguments. No backport needed, except if anonymization mechanism is backported.	2022-09-29 10:53:14 +02:00
Erwan Le Goas	1caa5351e5	MINOR: cli: Add an anonymization on a missed element in 'show server state' Add HA_ANON_CLI to the srv->hostname when using 'show servers state'. It can contain sensitive information like 'www....com' No backport needed, except if anonymization mechanism is backported.	2022-09-29 10:53:14 +02:00
Erwan Le Goas	9ac3ccb03f	MINOR: cli: use hash_ipanon to anonymized address Replace HA_ANON_CLI by hash_ipanon to anonynmized address like anonymizing address in the configuration file. No backport needed, except if anonymization mechanism is backported.	2022-09-29 10:53:14 +02:00
Erwan Le Goas	5eef1588a1	MINOR: tools: modify hash_ipanon in order to use it in cli Add a parameter hasport to return a simple hash or ipstring when ipstring has no port. Doesn't hash if scramble is null. Add option PA_O_PORT_RESOLVE to str2sa_range. Add a case UNIX. Those modification permit to use hash_ipanon in cli section in order to dump the same anonymization of address in the configuration file and with CLI. No backport needed, except if anonymization mechanism is backported.	2022-09-29 10:53:14 +02:00
Erwan Le Goas	3f4ae6194e	MINOR: cli: remove error message with 'set anon on\|off' Removed the error message in 'set anon on\|off', it's more user friendly: users use 'set anon on' even if the mode is already activated, and the same for 'set anon off'. That allows users to write the command line in the anonymized mode they want without errors. No backport needed, except if anonymization mechanism is backported.	2022-09-29 10:53:14 +02:00
Erwan Le Goas	2a2e46ff20	MINOR: cli: Add anonymization on a missed element for 'show sess all' Add an anonymization for an element missed in the first merge for 'show sess all'. No backport needed, except if anonymization mechanism is backported.	2022-09-29 10:53:14 +02:00
Aurelien DARRAGON	7fdba0ae54	BUG/MINOR: hlua: fixing hlua_http_msg_insert_data behavior hlua_http_msg_insert_data() function is called upon HTTPMessage.insert() method from lua script. This function did not work properly for multiple reasons: - An incorrect argument check was performed and prevented the user from providing optional offset argument. - Input and output variables were inverted and offset was not handled properly. The same bug was also fixed in hlua_http_msg_del_data(), see: 'BUG/MINOR: hlua: fixing hlua_http_msg_del_data behavior' The function now behaves as described in the documentation. This could be backported to 2.6 and 2.5.	2022-09-28 18:43:25 +02:00
Aurelien DARRAGON	d7c71b03d8	BUG/MINOR: hlua: fixing hlua_http_msg_del_data behavior GH issue #1885 reported that HTTPMessage.remove() did not work as expected. It turns out that underlying hlua_http_msg_del_data() function was not working properly due to input / output inversion as well as incorrect user offset handling. This patch fixes it so that the behavior is the one described in the documentation. This could be backported to 2.6 and 2.5.	2022-09-28 18:43:19 +02:00
Christopher Faulet	c5daf2801a	Revert "BUG/MINOR: config: don't count trailing spaces as empty arg" This reverts commit `5529424ef1`. Since this patch, HAProxy crashes when the first line of the configuration file contains more than one parameter because, on the first call of parse_line(), the output line is not allocated. Thus elements in the arguments array may point on invalid memory area. It may be considered as a bug to reference invalid memory area and should be fixed. But for now, it is safer to revert this patch If the reverted commit is backported, this one must be backported too.	2022-09-28 18:40:50 +02:00
Erwan Le Goas	5529424ef1	BUG/MINOR: config: don't count trailing spaces as empty arg In parse_line(), spaces increment the arg count and it is incremented again on '#' or end of line, resulting in an extra empty arg at the end of arg's list. The visible effect is that the reported arg count is in excess of 1. It doesn't seem to affect regular function but specialized ones like anonymisation depends on this count. This may need to be backported to stable versions.	2022-09-28 15:16:29 +02:00
William Lallemand	3a374eaeeb	BUG/MINOR: ring: fix the size check in ring_make_from_area() Fix the size check in ring_make_from_area() which is checking the size of the pointer instead of the size of the structure. No backport needed, 2.7 only.	2022-09-27 14:31:37 +02:00
Christopher Faulet	eaabf06031	BUG/MEDIUM: resolvers: Remove aborted resolutions from query_ids tree To avoid any UAF when a resolution is released, a mechanism was added to abort a resolution and delayed the released at the end of the current execution path. This mechanism depends on an hard assumption: Any reference on an aborted resolution must be removed. So, when a resolution is aborted, it is removed from the resolver lists and inserted into a death row list. However, a resolution may still be referenced in the query_ids tree. It is the tree containing all resolutions with a pending request. Because aborted resolutions are released outside the resolvers lock, it is possible to release a resolution on a side while a query ansswer is received and processed on another one. Thus, it is still possible to have a UAF because of this bug. To fix the issue, when a resolution is aborted, it is removed from any list, but it is also removed from the query_ids tree. This patch should solve the issue #1862 and may be related to #1875. It must be backported as far as 2.2.	2022-09-27 11:18:17 +02:00
Christopher Faulet	3ab72c66a0	BUG/MEDIUM: stconn: Reset SE descriptor when we fail to create a stream If stream_new() fails after the frontend SC is attached, the underlying SE descriptor is not properly reset. Among other things, SE_FL_ORPHAN flag is not set again. Because of this error, a BUG_ON() is triggered when the mux stream on the frontend side is destroyed. Thus, now, when stream_new() fails, SE_FL_ORPHAN flag is set on the SE descriptor and its stream-connector is set to NULL. This patch should solve the issue #1880. It must be backported to 2.6.	2022-09-27 11:18:11 +02:00
Christopher Faulet	4cfc038cb1	BUG/MINOR: stream: Perform errors handling in right order in stream_new() The frontend SC is attached before the backend one is allocated. Thus an allocation error on backend SC must be handled before an error on the frontend SC. This patch must be backported to 2.6.	2022-09-27 10:56:51 +02:00
William Lallemand	0a0512f76d	MINOR: mworker/cli: the mcli_reload bind_conf only send the reload status Upon a reload with the master CLI, the FD of the master CLI session is received by the internal socketpair listener. This session is used to display the status of the reload and then will close.	2022-09-24 16:35:23 +02:00
William Lallemand	56f73b21a5	MINOR: mworker: stores the mcli_reload bind_conf Stores the mcli_reload bind_conf in order to identify it later.	2022-09-24 15:56:25 +02:00
William Lallemand	21623b5949	MINOR: mworker: mworker_cli_proxy_new_listener() returns a bind_conf mworker_cli_proxy_new_listener() now returns a bind_conf * or NULL upon failure.	2022-09-24 15:51:27 +02:00
William Lallemand	68192b2cdf	MINOR: mworker: store and shows loading status The environment variable HAPROXY_LOAD_SUCCESS stores "1" if it successfully load the configuration and started, "0" otherwise. The "_loadstatus" master CLI command displays either "Loading failure!\n" or "Loading success.\n"	2022-09-24 15:44:42 +02:00
William Lallemand	479cb3ed3a	MINOR: mworker/cli: replace close() by fd_delete() Replace the close() call in cli_parse_reload() by a fd_delete() since the FD is one present in the fdtab.	2022-09-23 10:28:56 +02:00
Aurelien DARRAGON	b12d169ea3	BUG/MINOR: hlua: fixing ambiguous sizeof in hlua_load_per_thread As pointed out by chipitsine in GH #1879, coverity complains about a sizeof with char ** type where it should be char . This was introduced in 'MINOR: hlua: Allow argument on lua-lod(-per-thread) directives' (`ae6b568`) Luckily this had no effect so far because on most platforms sizeof(char ) == sizeof(char ), but this can not be safely assumed for portability reasons. The fix simply changes the argument to sizeof so that it refers to '*per_thread_load[len]' instead of 'per_thread_load[len]'. No backport needed.	2022-09-23 09:50:11 +02:00
William Lallemand	ec059c249e	MEDIUM: mworker/cli: keep the connection of the FD that ask for a reload When using the "reload" command over the master CLI, all connections to the master CLI were cut, this was unfortunate because it could have been used to implement a synchronous reload command. This patch implements an architecture to keep the connection alive after the reload. The master CLI is now equipped with a listener which uses a socketpair, the 2 FDs of this socketpair are stored in the mworker_proc of the master, which the master keeps via the environment variable. ipc_fd[1] is used as a listener for the master CLI. During the "reload" command, the CLI will send the FD of the current session over ipc_fd[0], then the reload is achieved, so the master won't handle the recv of the FD. Once reloaded, ipc_fd[1] receives the FD of the session, so the connection is preserved. Of course it is a new context, so everything like the "prompt mode" are lost. Only the FD which performs the reload is kept.	2022-09-22 18:16:19 +02:00
Erwan Le Goas	d2605cf0e5	BUG/MINOR: anon: memory illegal accesses in tools.c with hash_anon and hash_ipanon chipitsine reported in github issue #1872 that in function hash_anon and hash_ipanon, index_hash can be equal to NB_L_HASH_WORD and can reach an inexisting line table, the table is initialized hash_word[NB_L_HASH_WORD][20]; so hash_word[NB_L_HASH_WORD] doesn't exist. No backport needed, except if anonymization mechanism is backported.	2022-09-22 15:44:13 +02:00
Thierry Fournier	ae6b56800f	MINOR: hlua: Allow argument on lua-lod(-per-thread) directives Allow per-lua file argument which makes multiples configuration easier to handle This patch fixes issue #1609.	2022-09-22 15:24:29 +02:00
Thierry Fournier	70e38e91b4	BUG/MINOR: hlua: Remove \n in Lua error message built with memprintf Because memprintf return an error to the caller and not on screen. the function which perform display of message on the right output is in charge of adding \n if it is necessary. This patch may be backported.	2022-09-21 16:02:40 +02:00
wrightlaw	9a8d8a3fd0	BUG/MINOR: smtpchk: SMTP Service check should gracefully close SMTP transaction At present option smtpchk closes the TCP connection abruptly on completion of service checking, even if successful. This can result in a very high volume of errors in backend SMTP server logs. This patch ensures an SMTP QUIT is sent and a positive 2xx response is received from the SMTP server prior to disconnection. This patch depends on the following one: * MINOR: smtpchk: Update expect rule to fully match replies to EHLO commands This patch should fix the issue #1812. It may be backported as far as 2.2 with the commit above On the 2.2, proxy_parse_smtpchk_opt() function is located in src/check.c [cf: I updated reg-tests script accordingly]	2022-09-21 16:01:42 +02:00
Christopher Faulet	2ec1ffaed0	MINOR: smtpchk: Update expect rule to fully match replies to EHLO commands The response to EHLO command is a multiline reply. However the corresponding expect rule only match on the first line. For now, it is not an issue. But to be able to send the QUIT command and gracefully close the connection, we must be sure to consume the full EHLO reply first. To do so, the regex has been updated to match all 2xx lines at a time.	2022-09-21 15:11:26 +02:00
Willy Tarreau	4eaf85f5d9	MINOR: clock: do not update the global date too often Tests with forced wakeups on a 24c/48t machine showed that we're caping at 7.3M loops/s, which means 6.6 microseconds of loop delay without having anything to do. This is caused by two factors: - the load and update of the now_offset variable - the update of the global_now variable What is happening is that threads are not running within the one- microsecond time precision provided by gettimeofday(), so each thread waking up sees a slightly different date and causes undesired updates to global_now. But worse, these undesired updates mean that we then have to adjust the now_offset to match that, and adds significant noise to this variable, which then needs to be updated upon each call. By only allowing sightly less precision we can completely eliminate that contention. Here we're ignoring the 5 lowest bits of the usec part, meaning that the global_now variable may be off by up to 31 us (16 on avg). The variable is only used to correct the time drift some threads might be observing in environments where CPU clocks are not synchronized, and it's used by freq counters. In both cases we don't need that level of precision and even one millisecond would be pretty fine. We're just 30 times better at almost no cost since the global_now and now_offset variables now only need to be updated 30000 times a second in the worst case, which is unnoticeable. After this change, the wakeup rate jumped from 7.3M/s to 66M/s, meaning that the loop delay went from 6.6us to 0.73us, that's a 9x improvement when under load! With real tasks we're seeing a boost from 28M to 52M wakeups/s. The clock_update_global_date() function now only takes 1.6%, it's good enough so that we don't need to go further.	2022-09-21 09:06:28 +02:00
Willy Tarreau	58b73f9fa8	MINOR: pollers: only update the local date during busy polling This patch modifies epoll, kqueue and evports (the 3 pollers that support busy polling) to only update the local date in the inner polling loop, the global one being done when leaving the loop. Testing with epoll on a 24c/48t machine showed a boost from 53M to 352M loops/s, indicating that the loop was spending 85% of its time updating the global date or causing side effects (which was confirmed with perf top showing 67% in clock_update_global_date() alone).	2022-09-21 09:06:28 +02:00
Willy Tarreau	a700420671	MINOR: clock: split local and global date updates Pollers that support busy polling spend a lot of time (and cause contention) updating the global date when they're looping over themselves while it serves no purpose: what's needed is only an update on the local date to know when to stop looping. This patch splits clock_pudate_date() into a pair of local and global update functions, so that pollers can be easily improved.	2022-09-21 09:06:28 +02:00
Aurelien DARRAGON	ae1e14d65b	CLEANUP: tools: removing escape_chunk() function Func is not used anymore. See e3bde807d.	2022-09-20 16:25:30 +02:00
Aurelien DARRAGON	c5bff8e550	BUG/MINOR: log: improper behavior when escaping log data Patrick Hemmer reported an improper log behavior when using log-format to escape log data (+E option): Some bytes were truncated from the output: - escape_string() function now takes an extra parameter that allow the caller to specify input string stop pointer in case the input string is not guaranteed to be zero-terminated. - Minors checks were added into lf_text_len() to make sure dst string will not overflow. - lf_text_len() now makes proper use of escape_string() function. This should be backported as far as 1.8.	2022-09-20 16:25:30 +02:00
Christopher Faulet	b0b8e9bbd2	BUG/MINOR: mux-h1: Account consumed output data on synchronous connection error The commit `372b38f935` ("BUG/MEDIUM: mux-h1: Handle connection error after a synchronous send") introduced a bug. In h1_snd_buf(), consumed data are not properly accounted if a connection error is detected. Indeed, data are consumed when the output buffer is filled. But, on connection error, we exit from the loop without incremented total variable accordingly. When this happens, this leaves the channel buffer in an inconsistent state. The buffer may be empty with some output at the channel level. Because an error is reported, it is harmless. But it is safer to fix this bug now to avoid any regression in future. This patch must be backported as far as 2.2.	2022-09-20 16:23:45 +02:00
Amaury Denoyelle	0ed617ac2f	BUG/MEDIUM: mux-quic: properly trim HTX buffer on snd_buf reset MUX QUIC snd_buf operation whill return early if a qcs instance is resetted. In this case, HTX is left untouched and the callback returns the whole bufer size. This lead to an undefined behavior as the stream layer is notified about a transfer but does not see its HTX buffer emptied. In the end, the transfer may stall which will lead to a leak on session. To fix this, HTX buffer is now resetted when snd_buf is short-circuited. This should fix the issue as now the stream layer can continue the transfer until its completion. This patch has already been tested by Tristan and is reported to solve the github issue #1801. This should be backported up to 2.6.	2022-09-20 15:35:33 +02:00
Amaury Denoyelle	9534e59bb9	MINOR: mux-quic: refactor snd_buf Factorize common code between h3 and hq-interop snd_buf operation. This is inserted in MUX QUIC snd_buf own callback. The h3/hq-interop API has been adjusted to directly receive a HTX message instead of a plain buf. This led to extracting part of MUX QUIC snd_buf in qmux_http module. This should be backported up to 2.6.	2022-09-20 15:35:29 +02:00
Amaury Denoyelle	d80fbcaca2	REORG: mux-quic: export HTTP related function in a dedicated file Extract function dealing with HTX outside of MUX QUIC. For the moment, only rcv_buf stream operation is concerned. The main objective is to be able to support both TCP and HTTP proxy mode with a common base and add specialized modules on top of it. This should be backported up to 2.6.	2022-09-20 15:35:23 +02:00
Amaury Denoyelle	36d50bff22	REORG: mux-quic: extract traces in a dedicated source file QUIC MUX implements several APIs to interface with stream, quic-conn and app-ops layers. It is planified to better separate this roles, possibly by using several files. The first step is to extract QUIC MUX traces in a dedicated source files. This will allow to reuse traces in multiple files. The main objective is to be able to support both TCP and HTTP proxy mode with a common base and add specialized modules on top of it. This should be backported up to 2.6.	2022-09-20 15:35:09 +02:00
Amaury Denoyelle	3dc4e5a5b9	BUG/MINOR: mux-quic: do not keep detached qcs with empty Tx buffers A qcs instance free may be postponed in stream detach operation if the stream is not locally closed. This condition is there to achieve transfering data still present in Tx buffer. Once all data have been emitted to quic-conn layer, qcs instance can be released. However, the stream is only closed locally if HTX EOM has been seen or it has been resetted. In case the transfer finished without EOM, a detached qcs won't be freed even if there is no more activity on it. This bug was not reproduced but was found on code analysis. Its precise impact is unknown but it should not cause any leak as all qcs instances are freed with its parent qcc connection : this should eventually happen on MUX timeout or QUIC idle timeout. To adjust this, condition to mark a stream as locally closed has been extended. On qcc_streams_sent_done() notification, if its Tx buffer has been fully transmitted, it will be closed if either FIN STREAM was set or the stream is detached. This must be backported up to 2.6.	2022-09-20 10:46:59 +02:00
Willy Tarreau	9f4f6b038c	OPTIM: hpack-huff: reduce the cache footprint of the huffman decoder Some tables are currently used to decode bit blocks and lengths. We do see such lookups in perf top. We have 4 512-byte tables and one 64-byte one. Looking closer, the second half of the table (length) has so few variations that most of the time it will be computed in a single "if", and never more than 3. This alone allows to cut the tables in half. In addition, one table (bits 15-11) is only 32-element long, while another one (bits 11-4) starts at 0x60, so we can merge the two as they do not overlap, and further save size. We're now down to 4 256-entries tables. This is visible in h3 and h2 where the max request rate is slightly higher (e.g. +1.6% for h2). The huff_dec() function got slightly larger but the overall code size shrunk: $ nm --size haproxy-before \| grep huff_dec 000000000000029e T huff_dec $ nm --size haproxy-after \| grep huff_dec 0000000000000345 T huff_dec $ size haproxy-before haproxy-after text data bss dec hex filename 7591126 569268 2761348 10921742 a6a70e haproxy-before 7591082 568180 2761348 10920610 a6a2a2 haproxy-after	2022-09-20 07:41:58 +02:00
Miroslav Zagorac	cbfee3a9f6	MINOR: httpclient: enabled the use of SNI presets This commit allows setting SNI outside http_client.c code.	2022-09-19 14:39:28 +02:00
Miroslav Zagorac	133e2a23d0	CLEANUP: httpclient: deleted unused variables The locally defined static variables 'httpclient_srv_raw' and 'httpclient_srv_ssl' are not used anywhere in the source code, except that they are set in the httpclient_precheck() function.	2022-09-19 14:39:28 +02:00
Amaury Denoyelle	afb7b9d8e5	BUG/MEDIUM: mux-quic: fix nb_hreq decrement nb_hreq is a counter on qcc for active HTTP requests. It is incremented for each qcs where a full HTTP request was received. It is decremented when the stream is closed locally : - on HTTP response fully transmitted - on stream reset A bug will occur if a stream is resetted without having processed a full HTTP request. nb_hreq will be decremented whereas it was not incremented. This will lead to a crash when building with DEBUG_STRICT=2. If BUG_ON_HOT are not active, nb_hreq counter will wrap which may break the timeout logic for the connection. This bug was triggered on haproxy.org. It can be reproduced by simulating the reception of a STOP_SENDING frame instead of a STREAM one by patching qc_handle_strm_frm() : + if (quic_stream_is_bidi(strm_frm->id)) + qcc_recv_stop_sending(qc->qcc, strm_frm->id, 0); + //ret = qcc_recv(qc->qcc, strm_frm->id, strm_frm->len, + // strm_frm->offset.key, strm_frm->fin, + // (char *)strm_frm->data); To fix this bug, a qcs is now flagged with a new QC_SF_HREQ_RECV. This is set when the full HTTP request is received. When the stream is closed locally, nb_hreq will be decremented only if this flag was set. This must be backported up to 2.6.	2022-09-19 12:12:21 +02:00
Erwan Le Goas	b0c0501516	MINOR: config: add command-line -dC to dump the configuration file This commit adds a new command line option -dC to dump the configuration file. An optional key may be appended to -dC in order to produce an anonymized dump using this key. The anonymizing process uses the same algorithm as the CLI so that the same key will produce the same hashes for the same identifiers. This way an admin may share an anonymized extract of a configuration to match against live dumps. Note that key 0 will not anonymize the output. However, in any case, the configuration is dumped after tokenizing, thus comments are lost.	2022-09-17 11:27:09 +02:00
Erwan Le Goas	acfdf7600b	MINOR: cli: anonymize 'show servers state' and 'show servers conn' Modify proxy.c in order to anonymize the following confidential data on commands 'show servers state' and 'show servers conn': - proxy name - server name - server address	2022-09-17 11:27:09 +02:00
Erwan Le Goas	57e35f4b87	MINOR: cli: anonymize commands 'show sess' and 'show sess all' Modify stream.c in order to hash the following confidential data if the anonymized mode is enabled: - configuration elements such as frontend/backend/server names - IP addresses	2022-09-17 11:27:09 +02:00
Erwan Le Goas	54966dffda	MINOR: anon: store the anonymizing key in the CLI's appctx In order to allow users to dump internal states using a specific key without changing the global one, we're introducing a key in the CLI's appctx. This key is preloaded from the global one when "set anon on" is used (and if none exists, a random one is assigned). And the key can optionally be assigned manually for the whole CLI session. A "show anon" command was also added to show the anon state, and the current key if the users has sufficient permissions. In addition, a "debug dev hash" command was added to test the feature.	2022-09-17 11:27:09 +02:00
Erwan Le Goas	fad9da83da	MINOR: anon: store the anonymizing key in the global structure Add a uint32_t key in global to hash words with it. A new CLI command 'set global-key <key>' was added to change the global anonymizing key. The global may also be set in the configuration using the global "anonkey" directive. For now this key is not used.	2022-09-17 11:24:53 +02:00
Erwan Le Goas	9c76637fff	MINOR: anon: add new macros and functions to anonymize contents These macros and functions will be used to anonymize strings by producing a short hash. This will allow to match config elements against dump elements without revealing the original data. This will later be used to anonymize configuration parts and CLI commands output. For now only string, identifiers and addresses are supported, but the model is easily extensible.	2022-09-17 11:24:53 +02:00
Willy Tarreau	85af760704	BUILD: fd: fix a build warning on the DWCAS Ilya reported in issue #1816 a build warning on armhf (promoted to error here since -Werror): src/fd.c: In function fd_rm_from_fd_list: src/fd.c:209:87: error: passing argument 3 of __ha_cas_dw discards volatile qualifier from pointer target type [-Werror=discarded-array-qualifiers] 209 \| unlikely(!_HA_ATOMIC_DWCAS(((long )&fdtab[fd].update), (uint32_t )&cur_list.u32, &next_list.u32)) \| ^~~~~~~~~~~~~~ This happens only on such an architecture because the DWCAS requires the pointer not the value, and gcc seems to be needlessly picky about reading a const from a volatile! This may safely be backported to older versions.	2022-09-17 11:20:44 +02:00
Willy Tarreau	da9f258759	BUG/MEDIUM: captures: free() an error capture out of the proxy lock Ed Hein reported in github issue #1856 some occasional watchdog panics in 2.4.18 showing extreme contention on the proxy's lock while the libc was in malloc()/free(). One cause of this problem is that we call free() under the proxy's lock in proxy_capture_error(), which makes no sense since if we can free the object under the lock after it's been detached, we can also free it after releasing the lock (since it's not referenced anymore). This should be backported to all relevant versions, likely all supported ones.	2022-09-17 11:07:19 +02:00
cui fliter	a94bedc0de	CLEANUP: quic,ssl: fix tiny typos in C comments This fixes 4 tiny and harmless typos in mux_quic.c, quic_tls.c and ssl_sock.c. Originally sent via GitHub PR #1843. Signed-off-by: cui fliter <imcusg@gmail.com> [Tim: Rephrased the commit message] [wt: further complete the commit message]	2022-09-17 10:59:59 +02:00
Aurelien DARRAGON	8d0ff28406	BUG/MEDIUM: server: segv when adding server with hostname from CLI When calling 'add server' with a hostname from the cli (runtime), str2sa_range() does not resolve hostname because it is purposely called without PA_O_RESOLVE flag. This leads to 'srv->addr_node.key' being NULL. According to Willy it is fine behavior, as long as we handle it properly, and is already handled like this in srv_set_addr_desc(). This patch fixes GH #1865 by adding an extra check before inserting 'srv->addr_node' into 'be->used_server_addr'. Insertion and removal will be skipped if 'addr_node.key' is NULL. It must be backported to 2.6 and 2.5 only.	2022-09-17 06:30:59 +02:00
Amaury Denoyelle	d1310f8d32	BUG/MINOR: mux-quic: do not remotely close stream too early A stream is considered as remotely closed once we have received all the data with the FIN bit set. The condition to close the stream was wrong. In particular, if we receive an empty STREAM frame with FIN bit set, this would have close the stream even if we do not have yet received all the data. The condition is now adjusted to ensure that Rx buffer contains all the data up to the stream final size. In most cases, this bug is harmless. However, if compiled with DEBUG_STRICT=2, a BUG_ON_HOT crash would have been triggered if close is done too early. This was most notably the case sometimes on interop test suite with quinn or kwik clients. This can also be artificially reproduced by simulating reception of an empty STREAM frame with FIN bit set in qc_handle_strm_frm() : + if (strm_frm->fin) { + qcc_recv(qc->qcc, strm_frm->id, 0, + strm_frm->len, strm_frm->fin, + (char )strm_frm->data); + } ret = qcc_recv(qc->qcc, strm_frm->id, strm_frm->len, strm_frm->offset.key, strm_frm->fin, (char )strm_frm->data); This must be backported up to 2.6.	2022-09-16 14:17:27 +02:00
Amaury Denoyelle	8d4ac48d3d	CLEANUP: mux-quic: remove stconn usage in h3/hq Small cleanup on snd_buf for application protocol layer. * do not export h3_snd_buf * replace stconn by a qcs argument. This is better as h3/hq-interop only uses the qcs instance. This should be backported up to 2.6.	2022-09-16 13:53:30 +02:00
Christopher Faulet	18ad15f5c4	REORG: mux-h1: extract flags and enums into mux_h1-t.h The same was performed for the H2 multiplexer. H1C and H1S flags are moved in a dedicated header file. It will be mainly used to be able to decode mux-h1 flags from the flags utility. In this patch, we only move the flags to mux_h1-t.h.	2022-09-15 11:01:59 +02:00
Amaury Denoyelle	f8aaf8bdfa	BUG/MEDIUM: mux-quic: fix crash on early app-ops release H3 SETTINGS emission has recently been delayed. The idea is to send it with the first STREAM to reduce sendto syscall invocation. This was implemented in the following patch : `3dd79d378c` MINOR: h3: Send the h3 settings with others streams (requests) This patch works fine under nominal conditions. However, it will cause a crash if a HTTP/3 connection is released before having sent any data, for example when receiving an invalid first request. In this case, qc_release will first free qcc.app_ops HTTP/3 application protocol layer via release callback. Then qc_send is called to emit any closing frames built by app_ops release invocation. However, in qc_send, as no data has been sent, it will try to complete application layer protocol intialization, with a SETTINGS emission for HTTP/3. Thus, qcc.app_ops is reused, which is invalid as it has been just freed. This will cause a crash with h3_finalize in the call stack. This bug can be reproduced artificially by generating incomplete HTTP/3 requests. This will in time trigger http-request timeout without any data send. This is done by editing qc_handle_strm_frm function. - ret = qcc_recv(qc->qcc, strm_frm->id, strm_frm->len, + ret = qcc_recv(qc->qcc, strm_frm->id, strm_frm->len - 1, strm_frm->offset.key, strm_frm->fin, (char *)strm_frm->data); To fix this, application layer closing API has been adjusted to be done in two-steps. A new shutdown callback is implemented : it is used by the HTTP/3 layer to generate GOAWAY frame in qc_release prologue. Application layer context qcc.app_ops is then freed later in qc_release via the release operation which is now only used to liberate app layer ressources. This fixes the problem as the intermediary qc_send invocation will be able to reuse app_ops before it is freed. This patch fixes the crash, but it would be better to adjust H3 SETTINGS emission in case of early connection closing : in this case, there is no need to send it. This should be implemented in a future patch. This should fix the crash recently experienced by Tristan in github issue #1801. This must be backported up to 2.6.	2022-09-15 10:41:44 +02:00
William Lallemand	95fc737fc6	MEDIUM: quic: separate path for rx and tx with set_encryption_secrets With quicTLS the set_encruption_secrets callback is always called with the read_secret and the write_secret. However this is not the case with libreSSL, which uses the set_read_secret()/set_write_secret() mecanism. It still provides the set_encryption_secrets() callback, which is called with a NULL parameter for the write_secret during the read, and for the read_secret during the write. The exchange key was not designed in haproxy to be called separately for read and write, so this patch allow calls with read or write key to NULL.	2022-09-14 18:16:37 +02:00
William Lallemand	992ad62e3c	MEDIUM: httpclient: allow to use another proxy httpclient_new_from_proxy() is a variant of httpclient_new() which allows to create the requests from a different proxy. The proxy and its 2 servers are now stored in the httpclient structure. The proxy must have been created with httpclient_create_proxy() to be used. The httpclient_postcheck() callback will finish the initialization of all proxies created with PR_CAP_HTTPCLIENT.	2022-09-13 17:12:38 +02:00
William Lallemand	54aec5f678	MEDIUM: httpclient: httpclient_create_proxy() creates a proxy for httpclient httpclient_create_proxy() is a function which creates a proxy that could be used for the httpclient. It will allocate a proxy, a raw server and an ssl server. This patch moves most of the code from httpclient_precheck() into a generic function httpclient_create_proxy(). The proxy will have the PR_CAP_HTTPCLIENT capability. This could be used for specifics httpclient instances that needs different proxy settings.	2022-09-13 17:12:38 +02:00
Emeric Brun	d6e581de4b	BUG/MEDIUM: sink: bad init sequence on tcp sink from a ring. The init of tcp sink, particularly for SSL, was done too early in the code, during parsing, and this can cause a crash specially if nbthread was not configured. This was detected by William using ASAN on a new regtest on log forward. This patch adds the 'struct proxy' created for a sink to a list and this list is now submitted to the same init code than the main proxies list or the log_forward's proxies list. Doing this, we are assured to use the right init sequence. It also removes the ini code for ssl from post section parsing. This patch should be backported as far as v2.2 Note: this fix uses 'goto' labels created by commit 'BUG/MAJOR: log-forward: Fix log-forward proxies not fully initialized' but this code didn't exist before v2.3 so this patch needs to be adapted for v2.2.	2022-09-13 17:03:30 +02:00
Willy Tarreau	6c0fadfb7d	REORG: mux-h2: extract flags and enums into mux_h2-t.h Originally in 1.8 we wanted to have an independent mux that could possibly be disabled and would not impose dependencies on the outside. Everything would fit into a single C file and that was fine. Nowadays muxes are unavoidable, and not being able to easily inspect them from outside is sometimes a bit of a pain. In particular, the flags utility still cannot be used to decode their flags. As a first step towards this, this patch moves the flags and enums to mux_h2-t.h, as well as the two state decoding inline functions. It also dropped the H2_SS_*_BIT defines that nobody uses. The mux_h2.c file remains the only one to include that for now.	2022-09-12 19:33:07 +02:00
Aurelien DARRAGON	a57786e87d	BUG/MINOR: listener: null pointer dereference suspected by coverity Please refer to GH #1859 for more info. Coverity suspected improper proxy pointer handling. Without the fix it is considered safe for the moment, but it might not be the case in the future as we want to keep the ability to have isolated listeners. Making sure stop_listener(), pause_listener(), resume_listener() and listener_release() functions make proper use of px pointer in that context. No need for backport except if multi-connection protocols (ie:FTP) were to be backported as well.	2022-09-12 10:12:18 +02:00
Aurelien DARRAGON	187396e34e	CLEANUP: listener: function comment typo in stop_listener() A minor typo related to stop_listener() function comment was introduced in `0013288`. This makes stop_listener() function comment easier to read.	2022-09-12 10:12:13 +02:00
Christopher Faulet	af5336fd23	BUG/MINOR: mux-h1: Increment open_streams counter when H1 stream is created Since this counter was added, it was incremented at the wrong place for client streams. It was incremented when the stream-connector (formely the conn-stream) was created while it should be done when the H1 stream is created. Thus, on parsing error, on H1>H2 upgrades or TCP>H1 upgrades, the counter is not incremented. However, it is always decremented when the H1 stream is destroyed. On bakcned side, there is no issue. This patch must be backported to 2.6.	2022-09-12 09:54:11 +02:00
Willy Tarreau	af985e0151	CLEANUP: pollers: remove dead code in the polling loop As reported by Ilya and Coverity in issue #1858, since recent commit `eea152ee6` ("BUG/MINOR: signals/poller: ensure wakeup from signals") which removed the test for the global signal flag from the pollers' loop, the remaining "wake" flag doesn't need to be tested since it already participates to zeroing the wait_time and will be caught on the previous line. Let's just remove that test now.	2022-09-12 09:35:44 +02:00
Aurelien DARRAGON	cddec0aef5	BUG/MINOR: stats: fixing stat shows disabled frontend status as 'OPEN' This patch adresses the issue #1626. Adding support for PR_FL_PAUSED flag in the function stats_fill_fe_stats(). The command 'show stat' now properly reports a disabled frontend using "PAUSED" state label. This patch depends on the following commits: - `7d00077fd5` "BUG/MEDIUM: proxy: ensure pause_proxy() and resume_proxy() own PROXY_LOCK". - `001328873c` "MINOR: listener: small API change" - `d46f437de6` "MINOR: proxy/listener: support for additional PAUSED state" It should be backported to 2.6, 2.5 and 2.4	2022-09-09 17:24:22 +02:00
Aurelien DARRAGON	d46f437de6	MINOR: proxy/listener: support for additional PAUSED state This patch is a prerequisite for #1626. Adding PAUSED state to the list of available proxy states. The flag is set when the proxy is paused at runtime (pause_listener()). It is cleared when the proxy is resumed (resume_listener()). It should be backported to 2.6, 2.5 and 2.4	2022-09-09 17:23:01 +02:00
Aurelien DARRAGON	001328873c	MINOR: listener: small API change A minor API change was performed in listener(.c/.h) to restore consistency between stop_listener() and (resume/pause)_listener() functions. LISTENER_LOCK was never locked prior to calling stop_listener(): lli variable hint is thus not useful anymore. Added PROXY_LOCK locking in (resume/pause)_listener() functions with related lpx variable hint (prerequisite for #1626). It should be backported to 2.6, 2.5 and 2.4	2022-09-09 17:23:01 +02:00
Aurelien DARRAGON	7d00077fd5	BUG/MEDIUM: proxy: ensure pause_proxy() and resume_proxy() own PROXY_LOCK There was a race involving hlua_proxy_* functions and some proxy management functions. pause_proxy() and resume_proxy() can be used directly from lua code, but that could lead to some race as lua code didn't make sure PROXY_LOCK was owned before calling the proxy functions. This patch makes sure it won't happen again elsewhere in the code by locking PROXY_LOCK directly in resume and pause proxy functions so that it's not the caller's responsibility anymore. (based on stop_proxy() behavior that was already safe prior to the patch) This should be backported to stable series. Note that the API will likely differ < 2.4	2022-09-09 17:23:01 +02:00
Matthias Wirth	eea152ee68	BUG/MINOR: signals/poller: ensure wakeup from signals Add self-wake in signal_handler() to fix a race condition with a signal coming in between checking signal_queue_len and entering polling sleep. The changes in commit `43c891dda` ("BUG/MINOR: signals/poller: set the poller timeout to 0 when there are signals") were insufficient. Move the signal_queue_len check from the poll implementations to run_poll_loop() to keep that logic in one place. The poll loops are terminated either by the parameter wake being set or wake up due to a write to their poller_wr_pipe by wake_thread() in signal_handler(). This fixes issue #1841. Must be backported in every stable version.	2022-09-09 11:15:22 +02:00
Frédéric Lécaille	3dd79d378c	MINOR: h3: Send the h3 settings with others streams (requests) This is the ->finalize application callback which prepares the unidirectional STREAM frames for h3 settings and wakeup the mux I/O handler to send them. As haproxy is at the same time always waiting for the client request, this makes haproxy call sendto() to send only about 20 bytes of stream data. Furthermore in case of heavy loss, this give less chances to short h3 requests to succeed. Drawback: as at this time the mux sends its streams by their IDs ascending order the stream 0 is always embedded before the unidirectional stream 3 for h3 settings. Nevertheless, as these settings may be lost and received after other h3 request streams, this is permitted by the RFC. Perhaps there is a better way to do. This will have to be checked with Amaury. Must be backported to 2.6.	2022-09-08 18:04:58 +02:00
Frédéric Lécaille	befcf7031d	MINOR: h3: Missing connection argument for a TRACE_LEAVE() argument This should help in debbuging issues to be able to associate this trace to a QUIC connection. Must be backported to 2.6.	2022-09-08 18:04:58 +02:00
Frédéric Lécaille	2eb5faa2ad	MINOR: h3: Add the quic_conn object to h3 traces This is very useful to associate h3 traces to a QUIC connection when debugging. Must be backported to 2.6.	2022-09-08 18:04:58 +02:00
Frédéric Lécaille	1c725aa9cd	BUG/MINOR: h3: Crash when h3 trace verbosity is "minimal" This was due to a missing check in h3_trace() about the first argument presence (connection) and h3_parse_settings_frm() which calls TRACE_LEAVE() without any argument. Then this argument was dereferenced. Must be backported to 2.6	2022-09-08 18:04:58 +02:00
Frédéric Lécaille	3c1b81fdd7	BUG/MINOR: quic: Trace fix about packet number space information. <qc> variable was confused with <qel>. The consequence was that it was always the same packet number space which was displayed: the first one (or the Initial packet number space). Must be backported to 2.6.	2022-09-08 18:04:58 +02:00
Frédéric Lécaille	bb995eafc7	BUG/MINOR: quic: Speed up the handshake completion only one time It is possible to speed up the handshake completion but only one time by connection as mentionned in RFC 9002 "6.2.3. Speeding up Handshake Completion". Add a flag to prevent this process to be run several times (see https://www.rfc-editor.org/rfc/rfc9002#name-speeding-up-handshake-compl). Must be backported to 2.6.	2022-09-08 18:04:58 +02:00
William Lallemand	43c891dda0	BUG/MINOR: signals/poller: set the poller timeout to 0 when there are signals When receiving a signal before entering the poller, and without any activity in the process, the poller will be entered with a timeout calculated without checking the signals. Since commit 4f59d3 ("MINOR: time: increase the minimum wakeup interval to 60s") the issue is much more visible because it could be stuck for 60s. When in mworker mode, if a worker quits and the SIGCHLD signal deliver at the right time to the master, this one could be stuck for the time of the timeout. This should fix issue #1841 Must be backported in every stable version.	2022-09-08 17:46:31 +02:00
Willy Tarreau	e86bc35672	MINOR: activity/cli: support sorting task profiling by total CPU time The new "bytime" sorting criterion uses the reported CPU time instead of the usage. This is convenient to spot tasks that are mostly reponsible for the CPU usage in a running process. It supports both the detailed and the aggregated format. The output looks like this: > show profiling tasks bytime Tasks activity: function calls cpu_tot cpu_avg lat_tot lat_avg qc_io_cb 117739 1.961m 999.1us 37.45s 318.1us <- h3_snd_buf@src/h3.c:1084 tasklet_wakeup process_stream 7376273 1.384m 11.26us 1.013h 494.2us <- stream_new@src/stream.c:563 task_wakeup process_stream 8104400 1.133m 8.389us 1.130h 502.0us <- sc_notify@src/stconn.c:1209 task_wakeup qc_io_cb 43280 45.76s 1.057ms 13.95s 322.3us <- qc_stream_desc_ack@src/quic_stream.c:128 tasklet_wakeup h1_io_cb 11025715 24.82s 2.251us 5.406m 29.42us <- sock_conn_iocb@src/sock.c:869 tasklet_wakeup quic_conn_app_io_cb 312861 23.86s 76.27us 2.373s 7.584us <- qc_lstnr_pkt_rcv@src/xprt_quic.c:6184 tasklet_wakeup_after qc_io_cb 37063 12.65s 341.4us 6.409s 172.9us <- qc_treat_acked_tx_frm@src/xprt_quic.c:1695 tasklet_wakeup h1_io_cb 4783520 11.79s 2.463us 1.419h 1.068ms <- conn_subscribe@src/connection.c:732 tasklet_wakeup sc_conn_io_cb 12269693 11.51s 938.0ns 2.117h 621.2us <- sc_app_chk_rcv_conn@src/stconn.c:762 tasklet_wakeup sc_conn_io_cb 6479006 10.94s 1.689us 7.984m 73.93us <- h1_wake_stream_for_recv@src/mux_h1.c:2600 tasklet_wakeup qc_io_cb 12011 10.72s 892.5us 2.120s 176.5us <- qcc_release_remote_stream@src/mux_quic.c:1200 tasklet_wakeup h2_io_cb 246423 6.225s 25.26us 56.52s 229.4us <- h2_snd_buf@src/mux_h2.c:6712 tasklet_wakeup h2_io_cb 137744 6.076s 44.11us 16.59s 120.4us <- sock_conn_iocb@src/sock.c:869 tasklet_wakeup quic_lstnr_dghdlr 323575 3.062s 9.462us 3.424m 634.9us <- quic_lstnr_dgram_dispatch@src/quic_sock.c:255 tasklet_wakeup sc_conn_io_cb 1206939 1.616s 1.338us 27.62m 1.373ms <- qcs_notify_send@src/mux_quic.c:529 tasklet_wakeup h2_io_cb 212370 251.2ms 1.182us 6.476s 30.49us <- h2c_restart_reading@src/mux_h2.c:856 tasklet_wakeup h1_io_cb 44109 197.0ms 4.466us 31.89s 723.0us <- h1_takeover@src/mux_h1.c:4085 tasklet_wakeup quic_conn_app_io_cb 3029 87.59ms 28.92us 999.0ms 329.8us <- qc_process_timer@src/xprt_quic.c:4635 tasklet_wakeup task_run_applet 40 35.77ms 894.3us 4.407ms 110.2us <- sc_applet_create@src/stconn.c:489 appctx_wakeup task_run_applet 18 27.36ms 1.520ms 19.56us 1.086us <- sc_app_chk_snd_applet@src/stconn.c:996 appctx_wakeup sc_conn_io_cb 2186 11.76ms 5.377us 963.0ms 440.5us <- h1_wake_stream_for_send@src/mux_h1.c:2610 tasklet_wakeup qc_io_cb 8 9.880ms 1.235ms 5.871ms 733.9us <- qcs_consume@src/mux_quic.c:800 tasklet_wakeup quic_conn_io_cb 4 5.951ms 1.488ms 38.85us 9.713us <- qc_lstnr_pkt_rcv@src/xprt_quic.c:6184 tasklet_wakeup_after qc_io_cb 101 4.975ms 49.26us 13.91ms 137.8us <- qc_process_timer@src/xprt_quic.c:4602 tasklet_wakeup h1_io_cb 2186 1.809ms 827.0ns 720.2ms 329.5us <- sock_conn_iocb@src/sock.c:849 tasklet_wakeup qc_process_timer 3031 1.735ms 572.0ns 1.153s 380.3us <- wake_expired_tasks@src/task.c:344 task_wakeup accept_queue_process 359 1.362ms 3.793us 80.32ms 223.7us <- listener_accept@src/listener.c:1099 tasklet_wakeup quic_conn_app_io_cb 2 921.1us 460.6us 203.1us 101.5us <- qc_xprt_start@src/xprt_quic.c:7122 tasklet_wakeup h1_timeout_task 2618 526.8us 201.0ns 1.121s 428.4us <- h1_release@src/mux_h1.c:1087 task_wakeup process_resolvers 316 283.3us 896.0ns 14.96ms 47.33us <- wake_expired_tasks@src/task.c:429 task_drop_running sc_conn_io_cb 420 235.6us 560.0ns 116.7ms 277.8us <- h2s_notify_recv@src/mux_h2.c:1298 tasklet_wakeup qc_idle_timer_task 1 225.5us 225.5us 506.0ns 506.0ns <- wake_expired_tasks@src/task.c:344 task_wakeup accept_queue_process 36 153.0us 4.250us 5.834ms 162.1us <- accept_queue_process@src/listener.c:165 tasklet_wakeup sc_conn_io_cb 18 54.05us 3.003us 11.50us 638.0ns <- sock_conn_iocb@src/sock.c:869 tasklet_wakeup h2_io_cb 6 38.88us 6.480us 2.089ms 348.2us <- h2_do_shutw@src/mux_h2.c:4656 tasklet_wakeup srv_cleanup_idle_conns 54 37.72us 698.0ns 14.21ms 263.1us <- wake_expired_tasks@src/task.c:429 task_drop_running sc_conn_io_cb 50 32.86us 657.0ns 28.83ms 576.5us <- qcs_notify_recv@src/mux_quic.c:519 tasklet_wakeup qc_io_cb 2 30.25us 15.12us 6.093us 3.046us <- qc_init@src/mux_quic.c:2057 tasklet_wakeup srv_cleanup_toremove_conns 1 27.16us 27.16us 905.6us 905.6us <- srv_cleanup_idle_conns@src/server.c:5948 task_wakeup task_run_applet 39 19.61us 502.0ns 818.7us 20.99us <- run_tasks_from_lists@src/task.c:652 task_drop_running quic_accept_run 2 15.46us 7.727us 305.5us 152.8us <- quic_accept_push_qc@src/quic_sock.c:458 tasklet_wakeup h2_timeout_task 32 12.91us 403.0ns 4.207ms 131.5us <- h2_release@src/mux_h2.c:1191 task_wakeup quic_conn_app_io_cb 1 9.645us 9.645us 1.445us 1.445us <- qc_process_timer@src/xprt_quic.c:4589 tasklet_wakeup > show profiling tasks bytime aggr Tasks activity: function calls cpu_tot cpu_avg lat_tot lat_avg qc_io_cb 212301 3.147m 889.5us 1.009m 285.2us process_stream 15503573 2.519m 9.747us 2.148h 498.7us h1_io_cb 15916733 36.95s 2.321us 1.535h 347.1us quic_conn_app_io_cb 318845 24.21s 75.92us 3.410s 10.70us sc_conn_io_cb 20037058 24.19s 1.207us 2.737h 491.8us h2_io_cb 596543 12.55s 21.04us 1.326m 133.4us quic_lstnr_dghdlr 326624 3.094s 9.473us 3.462m 635.9us task_run_applet 100 64.43ms 644.3us 5.285ms 52.85us quic_conn_io_cb 4 5.951ms 1.488ms 38.85us 9.713us qc_process_timer 3061 1.750ms 571.0ns 1.162s 379.5us accept_queue_process 396 1.521ms 3.840us 86.16ms 217.6us h1_timeout_task 2618 526.8us 201.0ns 1.121s 428.4us process_resolvers 319 286.0us 896.0ns 16.82ms 52.73us qc_idle_timer_task 1 225.5us 225.5us 506.0ns 506.0ns srv_cleanup_idle_conns 54 37.72us 698.0ns 14.21ms 263.1us srv_cleanup_toremove_conns 1 27.16us 27.16us 905.6us 905.6us quic_accept_run 2 15.46us 7.727us 305.5us 152.8us h2_timeout_task 32 12.91us 403.0ns 4.207ms 131.5us	2022-09-08 16:38:10 +02:00
Willy Tarreau	dc89b1806c	MINOR: activity/cli: support aggregating task profiling outputs By default we now dump stats between caller and callee, but by specifying "aggr" on the command line, stats get aggregated by callee again as it used to be before the feature was available. It may sometimes be helpful when comparing total call counts, though that's about all.	2022-09-08 16:32:17 +02:00
Willy Tarreau	64435aaa85	MINOR: tasks/activity: improve the caller-callee activity hash The previous dump already showed that the "other" category was getting a few entries. Let's proceed like for the memory profiling, by scanning a limited range of adjacent slots to find a spare one (16 max). That's pretty fast since close and likely prefetched and the comparison is cheap. The new dump now shows up to 45 entries below without "other": Now: Tasks activity: function calls cpu_tot cpu_avg lat_tot lat_avg task_run_applet 22 34.56ms 1.571ms 1.145ms 52.04us <- sc_applet_create@src/stconn.c:489 appctx_wakeup task_run_applet 21 11.11us 529.0ns 2.590ms 123.3us <- run_tasks_from_lists@src/task.c:652 task_drop_running task_run_applet 5 7.715ms 1.543ms 2.186us 437.0ns <- sc_app_chk_snd_applet@src/stconn.c:996 appctx_wakeup accept_queue_process 345 3.129ms 9.068us 72.84ms 211.1us <- listener_accept@src/listener.c:1099 tasklet_wakeup accept_queue_process 32 113.0us 3.529us 3.070ms 95.94us <- accept_queue_process@src/listener.c:165 tasklet_wakeup sc_conn_io_cb 5026032 3.037s 604.0ns 17.47m 208.5us <- sc_app_chk_rcv_conn@src/stconn.c:762 tasklet_wakeup sc_conn_io_cb 4361192 7.626s 1.748us 3.179m 43.74us <- h1_wake_stream_for_recv@src/mux_h1.c:2600 tasklet_wakeup sc_conn_io_cb 178293 275.4ms 1.544us 2.740m 922.0us <- qcs_notify_send@src/mux_quic.c:529 tasklet_wakeup sc_conn_io_cb 2561 15.84ms 6.185us 1.036s 404.4us <- h1_wake_stream_for_send@src/mux_h1.c:2610 tasklet_wakeup sc_conn_io_cb 453 261.4us 577.0ns 86.79ms 191.6us <- h2s_notify_recv@src/mux_h2.c:1298 tasklet_wakeup sc_conn_io_cb 89 50.05us 562.0ns 100.7ms 1.131ms <- qcs_notify_recv@src/mux_quic.c:519 tasklet_wakeup sc_conn_io_cb 8 19.04us 2.379us 472.5us 59.06us <- sock_conn_iocb@src/sock.c:869 tasklet_wakeup process_resolvers 50 57.50us 1.149us 1.116ms 22.32us <- wake_expired_tasks@src/task.c:429 task_drop_running srv_cleanup_idle_conns 8 5.669us 708.0ns 216.6us 27.08us <- wake_expired_tasks@src/task.c:429 task_drop_running process_stream 4599847 48.79s 10.61us 16.92m 220.7us <- sc_notify@src/stconn.c:1209 task_wakeup process_stream 4530081 52.82s 11.66us 14.92m 197.6us <- stream_new@src/stream.c:563 task_wakeup process_stream 15 201.7us 13.45us 31.58ms 2.105ms <- sc_app_chk_snd_conn@src/stconn.c:857 task_wakeup h1_io_cb 7861205 18.22s 2.317us 2.408m 18.38us <- sock_conn_iocb@src/sock.c:869 tasklet_wakeup h1_io_cb 474763 1.379s 2.905us 6.578m 831.4us <- conn_subscribe@src/connection.c:732 tasklet_wakeup h1_io_cb 34830 38.64ms 1.109us 18.85s 541.2us <- h1_takeover@src/mux_h1.c:4085 tasklet_wakeup h1_io_cb 2561 2.150ms 839.0ns 674.4ms 263.3us <- sock_conn_iocb@src/sock.c:849 tasklet_wakeup h1_timeout_task 2634 588.5us 223.0ns 890.5ms 338.1us <- h1_release@src/mux_h1.c:1087 task_wakeup h2_timeout_task 16 7.519us 469.0ns 1.146ms 71.63us <- h2_release@src/mux_h2.c:1191 task_wakeup h2_io_cb 99601 2.212s 22.21us 19.33s 194.1us <- h2_snd_buf@src/mux_h2.c:6712 tasklet_wakeup h2_io_cb 79777 146.6ms 1.837us 3.529s 44.24us <- h2c_restart_reading@src/mux_h2.c:856 tasklet_wakeup h2_io_cb 60698 2.259s 37.21us 4.704s 77.50us <- sock_conn_iocb@src/sock.c:869 tasklet_wakeup h2_io_cb 5 36.90us 7.380us 2.045ms 409.0us <- h2_do_shutw@src/mux_h2.c:4656 tasklet_wakeup qc_io_cb 26595 8.007s 301.1us 4.261s 160.2us <- qc_treat_acked_tx_frm@src/xprt_quic.c:1695 tasklet_wakeup qc_io_cb 7921 5.284s 667.1us 2.171s 274.1us <- qc_stream_desc_ack@src/quic_stream.c:128 tasklet_wakeup qc_io_cb 6229 5.851s 939.3us 1.856s 297.9us <- h3_snd_buf@src/h3.c:1084 tasklet_wakeup qc_io_cb 994 699.1ms 703.3us 174.9ms 176.0us <- qcc_release_remote_stream@src/mux_quic.c:1200 tasklet_wakeup qc_io_cb 65 9.883ms 152.0us 13.33ms 205.1us <- qc_process_timer@src/xprt_quic.c:4602 tasklet_wakeup qc_io_cb 1 293.5us 293.5us 105.9us 105.9us <- qcs_consume@src/mux_quic.c:800 tasklet_wakeup qc_io_cb 1 10.87us 10.87us 3.307us 3.307us <- qc_init@src/mux_quic.c:2057 tasklet_wakeup quic_conn_io_cb 2 2.531ms 1.265ms 2.839us 1.419us <- qc_lstnr_pkt_rcv@src/xprt_quic.c:6184 tasklet_wakeup_after quic_conn_app_io_cb 61392 2.620s 42.67us 268.0ms 4.365us <- qc_lstnr_pkt_rcv@src/xprt_quic.c:6184 tasklet_wakeup_after quic_conn_app_io_cb 408 10.56ms 25.88us 124.0ms 303.8us <- qc_process_timer@src/xprt_quic.c:4635 tasklet_wakeup quic_conn_app_io_cb 2 15.61us 7.806us 103.2us 51.59us <- qc_process_timer@src/xprt_quic.c:4589 tasklet_wakeup quic_conn_app_io_cb 1 410.6us 410.6us 11.52us 11.52us <- qc_xprt_start@src/xprt_quic.c:7122 tasklet_wakeup quic_lstnr_dghdlr 62716 409.2ms 6.523us 21.81s 347.8us <- quic_lstnr_dgram_dispatch@src/quic_sock.c:255 tasklet_wakeup qc_process_timer 410 245.4us 598.0ns 238.5ms 581.7us <- wake_expired_tasks@src/task.c:344 task_wakeup quic_accept_run 1 7.711us 7.711us 82.28us 82.28us <- quic_accept_push_qc@src/quic_sock.c:458 tasklet_wakeup	2022-09-08 16:25:36 +02:00
Willy Tarreau	3d4cdb198c	MEDIUM: tasks/activity: combine the called function with the caller Now instead of getting aggregate stats per called function, we have them per function AND per call place. The "byaddr" sort considers the function pointer first, then the call count, so that dominant callers of a given callee are instantly spotted. This allows to get sorted outputs like this: Tasks activity: function calls cpu_tot cpu_avg lat_tot lat_avg h1_io_cb 17357952 40.91s 2.357us 4.849m 16.76us <- sock_conn_iocb@src/sock.c:869 tasklet_wakeup sc_conn_io_cb 10357182 6.297s 607.0ns 27.93m 161.8us <- sc_app_chk_rcv_conn@src/stconn.c:762 tasklet_wakeup process_stream 9891131 1.809m 10.97us 53.61m 325.2us <- sc_notify@src/stconn.c:1209 task_wakeup process_stream 9823934 1.887m 11.52us 48.31m 295.1us <- stream_new@src/stream.c:563 task_wakeup sc_conn_io_cb 9347863 16.59s 1.774us 6.143m 39.43us <- h1_wake_stream_for_recv@src/mux_h1.c:2600 tasklet_wakeup h1_io_cb 501344 1.848s 3.686us 6.544m 783.2us <- conn_subscribe@src/connection.c:732 tasklet_wakeup sc_conn_io_cb 239717 492.3ms 2.053us 3.213m 804.3us <- qcs_notify_send@src/mux_quic.c:529 tasklet_wakeup h2_io_cb 173019 4.204s 24.30us 40.95s 236.7us <- h2_snd_buf@src/mux_h2.c:6712 tasklet_wakeup h2_io_cb 149487 424.3ms 2.838us 14.63s 97.87us <- h2c_restart_reading@src/mux_h2.c:856 tasklet_wakeup other 101893 4.626s 45.40us 14.84s 145.7us quic_lstnr_dghdlr 94389 614.0ms 6.504us 30.54s 323.6us <- quic_lstnr_dgram_dispatch@src/quic_sock.c:255 tasklet_wakeup quic_conn_app_io_cb 92205 3.735s 40.51us 390.9ms 4.239us <- qc_lstnr_pkt_rcv@src/xprt_quic.c:6184 tasklet_wakeup_after qc_io_cb 50355 19.01s 377.5us 10.65s 211.4us <- qc_treat_acked_tx_frm@src/xprt_quic.c:1695 tasklet_wakeup h1_io_cb 44427 155.0ms 3.489us 21.50s 484.0us <- h1_takeover@src/mux_h1.c:4085 tasklet_wakeup qc_io_cb 9018 4.924s 546.0us 3.084s 342.0us <- qc_stream_desc_ack@src/quic_stream.c:128 tasklet_wakeup h1_timeout_task 3236 1.172ms 362.0ns 1.119s 345.9us <- h1_release@src/mux_h1.c:1087 task_wakeup h1_io_cb 2804 7.974ms 2.843us 1.980s 706.0us <- sock_conn_iocb@src/sock.c:849 tasklet_wakeup sc_conn_io_cb 2804 33.44ms 11.92us 2.597s 926.2us <- h1_wake_stream_for_send@src/mux_h1.c:2610 tasklet_wakeup qc_io_cb 2623 2.669s 1.017ms 1.347s 513.5us <- h3_snd_buf@src/h3.c:1084 tasklet_wakeup qc_process_timer 662 526.4us 795.0ns 1.081s 1.633ms <- wake_expired_tasks@src/task.c:344 task_wakeup quic_conn_app_io_cb 648 12.62ms 19.47us 225.7ms 348.2us <- qc_process_timer@src/xprt_quic.c:4635 tasklet_wakeup accept_queue_process 286 1.571ms 5.494us 72.55ms 253.7us <- listener_accept@src/listener.c:1099 tasklet_wakeup process_resolvers 176 157.8us 896.0ns 7.835ms 44.52us <- wake_expired_tasks@src/task.c:429 task_drop_running qc_io_cb 167 10.71ms 64.12us 32.47ms 194.4us <- qc_process_timer@src/xprt_quic.c:4602 tasklet_wakeup sc_conn_io_cb 123 80.05us 650.0ns 50.35ms 409.4us <- qcs_notify_recv@src/mux_quic.c:519 tasklet_wakeup h2_timeout_task 32 30.69us 958.0ns 9.038ms 282.4us <- h2_release@src/mux_h2.c:1191 task_wakeup task_run_applet 24 33.79ms 1.408ms 5.838ms 243.3us <- sc_applet_create@src/stconn.c:489 appctx_wakeup accept_queue_process 17 56.34us 3.314us 7.505ms 441.5us <- accept_queue_process@src/listener.c:165 tasklet_wakeup srv_cleanup_toremove_conns 16 1.133ms 70.81us 5.685ms 355.3us <- srv_cleanup_idle_conns@src/server.c:5948 task_wakeup srv_cleanup_idle_conns 16 74.57us 4.660us 2.797ms 174.8us <- wake_expired_tasks@src/task.c:429 task_drop_running quic_conn_app_io_cb 12 786.9us 65.58us 2.042ms 170.1us <- qc_process_timer@src/xprt_quic.c:4589 tasklet_wakeup sc_conn_io_cb 9 20.55us 2.283us 2.475ms 275.0us <- sock_conn_iocb@src/sock.c:869 tasklet_wakeup h2_io_cb 8 34.12us 4.265us 1.784ms 223.0us <- h2_do_shutw@src/mux_h2.c:4656 tasklet_wakeup task_run_applet 4 6.615ms 1.654ms 2.306us 576.0ns <- sc_app_chk_snd_applet@src/stconn.c:996 appctx_wakeup quic_conn_io_cb 4 4.278ms 1.069ms 6.469us 1.617us <- qc_lstnr_pkt_rcv@src/xprt_quic.c:6184 tasklet_wakeup_after qc_io_cb 2 20.81us 10.40us 4.943us 2.471us <- qc_init@src/mux_quic.c:2057 tasklet_wakeup quic_conn_app_io_cb 2 752.9us 376.4us 63.97us 31.99us <- qc_xprt_start@src/xprt_quic.c:7122 tasklet_wakeup quic_accept_run 2 13.84us 6.920us 172.8us 86.42us <- quic_accept_push_qc@src/quic_sock.c:458 tasklet_wakeup qc_idle_timer_task 2 295.0us 147.5us 8.761us 4.380us <- wake_expired_tasks@src/task.c:344 task_wakeup qc_io_cb 1 867.1us 867.1us 812.8us 812.8us <- qcs_consume@src/mux_quic.c:800 tasklet_wakeup ... and calls sorted by address like this: Tasks activity: function calls cpu_tot cpu_avg lat_tot lat_avg task_run_applet 23 32.73ms 1.423ms 5.837ms 253.8us <- sc_applet_create@src/stconn.c:489 appctx_wakeup task_run_applet 4 6.615ms 1.654ms 2.306us 576.0ns <- sc_app_chk_snd_applet@src/stconn.c:996 appctx_wakeup accept_queue_process 285 1.566ms 5.495us 72.49ms 254.3us <- listener_accept@src/listener.c:1099 tasklet_wakeup accept_queue_process 17 56.34us 3.314us 7.505ms 441.5us <- accept_queue_process@src/listener.c:165 tasklet_wakeup sc_conn_io_cb 10357182 6.297s 607.0ns 27.93m 161.8us <- sc_app_chk_rcv_conn@src/stconn.c:762 tasklet_wakeup sc_conn_io_cb 9347863 16.59s 1.774us 6.143m 39.43us <- h1_wake_stream_for_recv@src/mux_h1.c:2600 tasklet_wakeup sc_conn_io_cb 239717 492.3ms 2.053us 3.213m 804.3us <- qcs_notify_send@src/mux_quic.c:529 tasklet_wakeup sc_conn_io_cb 2804 33.44ms 11.92us 2.597s 926.2us <- h1_wake_stream_for_send@src/mux_h1.c:2610 tasklet_wakeup sc_conn_io_cb 123 80.05us 650.0ns 50.35ms 409.4us <- qcs_notify_recv@src/mux_quic.c:519 tasklet_wakeup sc_conn_io_cb 9 20.55us 2.283us 2.475ms 275.0us <- sock_conn_iocb@src/sock.c:869 tasklet_wakeup process_resolvers 159 145.9us 917.0ns 7.823ms 49.20us <- wake_expired_tasks@src/task.c:429 task_drop_running srv_cleanup_idle_conns 16 74.57us 4.660us 2.797ms 174.8us <- wake_expired_tasks@src/task.c:429 task_drop_running srv_cleanup_toremove_conns 16 1.133ms 70.81us 5.685ms 355.3us <- srv_cleanup_idle_conns@src/server.c:5948 task_wakeup process_stream 9891130 1.809m 10.97us 53.61m 325.2us <- sc_notify@src/stconn.c:1209 task_wakeup process_stream 9823933 1.887m 11.52us 48.31m 295.1us <- stream_new@src/stream.c:563 task_wakeup h1_io_cb 17357952 40.91s 2.357us 4.849m 16.76us <- sock_conn_iocb@src/sock.c:869 tasklet_wakeup h1_io_cb 501344 1.848s 3.686us 6.544m 783.2us <- conn_subscribe@src/connection.c:732 tasklet_wakeup h1_io_cb 44427 155.0ms 3.489us 21.50s 484.0us <- h1_takeover@src/mux_h1.c:4085 tasklet_wakeup h1_io_cb 2804 7.974ms 2.843us 1.980s 706.0us <- sock_conn_iocb@src/sock.c:849 tasklet_wakeup h1_timeout_task 3236 1.172ms 362.0ns 1.119s 345.9us <- h1_release@src/mux_h1.c:1087 task_wakeup h2_timeout_task 32 30.69us 958.0ns 9.038ms 282.4us <- h2_release@src/mux_h2.c:1191 task_wakeup h2_io_cb 173019 4.204s 24.30us 40.95s 236.7us <- h2_snd_buf@src/mux_h2.c:6712 tasklet_wakeup h2_io_cb 149487 424.3ms 2.838us 14.63s 97.87us <- h2c_restart_reading@src/mux_h2.c:856 tasklet_wakeup h2_io_cb 8 34.12us 4.265us 1.784ms 223.0us <- h2_do_shutw@src/mux_h2.c:4656 tasklet_wakeup qc_io_cb 50355 19.01s 377.5us 10.65s 211.4us <- qc_treat_acked_tx_frm@src/xprt_quic.c:1695 tasklet_wakeup qc_io_cb 9018 4.924s 546.0us 3.084s 342.0us <- qc_stream_desc_ack@src/quic_stream.c:128 tasklet_wakeup qc_io_cb 2623 2.669s 1.017ms 1.347s 513.5us <- h3_snd_buf@src/h3.c:1084 tasklet_wakeup qc_io_cb 167 10.71ms 64.12us 32.47ms 194.4us <- qc_process_timer@src/xprt_quic.c:4602 tasklet_wakeup qc_io_cb 2 20.81us 10.40us 4.943us 2.471us <- qc_init@src/mux_quic.c:2057 tasklet_wakeup qc_io_cb 1 867.1us 867.1us 812.8us 812.8us <- qcs_consume@src/mux_quic.c:800 tasklet_wakeup qc_idle_timer_task 2 295.0us 147.5us 8.761us 4.380us <- wake_expired_tasks@src/task.c:344 task_wakeup quic_conn_io_cb 4 4.278ms 1.069ms 6.469us 1.617us <- qc_lstnr_pkt_rcv@src/xprt_quic.c:6184 tasklet_wakeup_after quic_conn_app_io_cb 92205 3.735s 40.51us 390.9ms 4.239us <- qc_lstnr_pkt_rcv@src/xprt_quic.c:6184 tasklet_wakeup_after quic_conn_app_io_cb 648 12.62ms 19.47us 225.7ms 348.2us <- qc_process_timer@src/xprt_quic.c:4635 tasklet_wakeup quic_conn_app_io_cb 12 786.9us 65.58us 2.042ms 170.1us <- qc_process_timer@src/xprt_quic.c:4589 tasklet_wakeup quic_conn_app_io_cb 2 752.9us 376.4us 63.97us 31.99us <- qc_xprt_start@src/xprt_quic.c:7122 tasklet_wakeup quic_lstnr_dghdlr 94389 614.0ms 6.504us 30.54s 323.6us <- quic_lstnr_dgram_dispatch@src/quic_sock.c:255 tasklet_wakeup qc_process_timer 662 526.4us 795.0ns 1.081s 1.633ms <- wake_expired_tasks@src/task.c:344 task_wakeup quic_accept_run 2 13.84us 6.920us 172.8us 86.42us <- quic_accept_push_qc@src/quic_sock.c:458 tasklet_wakeup other 101892 4.626s 45.40us 14.84s 145.7us It already becomes visible that some tasks have different very costs depending where they're called (e.g. process_stream). The method used to wake them up is also shown. Applets are handled specially and shown as appctx_wakeup.	2022-09-08 16:21:22 +02:00
Willy Tarreau	41e701e2c1	DEBUG: quic: export the few task handlers that often appear in task dumps The following task/tasklet handlers often appear in "show profiling tasks" but were not resolved since static: qc_io_cb, quic_conn_app_io_cb, process_timer, quic_accept_run, qc_idle_timer_task This commit simply exports them so they can be resolved now. "process_timer" which was a bit too generic and renamed to qc_process_timer.	2022-09-08 16:13:38 +02:00
Willy Tarreau	0fbc16cfb9	DEBUG: resolvers: unstatify process_resolvers() to make it appear in profiling The function appears like this in "show profiling tasks", so let's export it: function calls cpu_tot cpu_avg lat_tot lat_avg main+0x1463f0 92 77.28us 839.0ns 2.018ms 21.93us <- wake_expired_tasks@src/task.c:429 task_drop_running	2022-09-08 16:13:38 +02:00
Willy Tarreau	a3423873fe	CLEANUP: activity: make the number of sched activity entries more configurable This removes all the hard-coded 8-bit and 256 entries to use a pair of macros instead so that we can more easily experiment with larger table sizes if needed.	2022-09-08 14:55:09 +02:00
Willy Tarreau	a9a2384612	CLEANUP: sched: remove duplicate code in run_tasks_from_list() Now that ->wake_date is common to tasks and tasklets, we don't need anymore to carry a duplicate control block to read and update it for tasks and tasklets. And given that this code was present early in the if/else fork between tasks and tasklets, taking it out of the block allows to move the task part into a more visible "else" branch that also allows to factor the epilogue that resets th_ctx->current and updates profile_entry->cpu_time, which also used to be duplicated. Overall, doing just that saved 253 bytes in the function, or ~1/6, which is not bad considering that it's on a hot path. And the code got much ore readable.	2022-09-08 14:30:38 +02:00
Willy Tarreau	d96d214b4c	CLEANUP: debug: use struct ha_caller for memstat The memstats code currently defines its own file/function/line number, type and extra pointer. We don't need to keep them separate and we can easily replace them all with just a struct ha_caller. Note that the extra pointer could be converted to a pool ID stored into arg8 or arg32 and be dropped as well, but this would first require to define IDs for pools (which we currently do not have).	2022-09-08 14:19:15 +02:00
Willy Tarreau	4c1bc01f31	CLEANUP: activity: make taskprof use ptr_hash() There's no more point using a different hash function here, xxh64 is of course better distributed but we really don't care so let's unify the code.	2022-09-08 14:19:15 +02:00
Willy Tarreau	245d32fe8f	CLEANUP: activity: make memprof use the generic ptr_hash() function There's no need to keep a local version of that function anymore.	2022-09-08 14:19:15 +02:00
Willy Tarreau	6a28a30efa	MINOR: tasks: do not keep cpu and latency times in struct task It was a mistake to put these two fields in the struct task. This was added in 1.9 via commit `9efd7456e` ("MEDIUM: tasks: collect per-task CPU time and latency"). These fields are used solely by streams in order to report the measurements via the lat_ns* and cpu_ns* sample fetch functions when task profiling is enabled. For the rest of the tasks, this is pure CPU waste when profiling is enabled, and memory waste 100% of the time, as the point where these latencies and usages are measured is in the profiling array. Let's move the fields to the stream instead, and have process_stream() retrieve the relevant info from the thread's context. The struct task is now back to 120 bytes, i.e. almost two cache lines, with 32 bit still available.	2022-09-08 14:19:15 +02:00
Willy Tarreau	beee600491	BUG/MINOR: stream/sched: take into account CPU profiling for the last call When task profiling is enabled, the reported CPU time for short requests and responses (e.g. redirect) is always zero in the logs, because process_stream() is only called once and the CPU time is measured after it returns. This is particuarly annoying when dealing with denies and in general anything that deals with parasitic traffic because it can be difficult to figure where the CPU is spent. The solution taken in this patch consists in having process_stream() update the cpu time itself before logging and quitting. It's very simple. It will not take into account the time taken to produce the log nor freeing the stream, but that's marginal compared to always logging zero. The task's wake_date is also reset so that the scheduler doesn't have to perform these operations again. This is dependent on the following patch: MINOR: sched: store the current profile entry in the thread context It should be backported to 2.6 as it does help for troubleshooting.	2022-09-08 14:19:15 +02:00
Willy Tarreau	1efddfa6bf	MINOR: sched: store the current profile entry in the thread context The profile entry that corresponds to the current task/tasklet being profiled is now stored into the thread's context. This will allow it to be accessed from the tasks themselves. This is needed for an upcoming fix.	2022-09-08 14:19:15 +02:00
Willy Tarreau	62b5b96bcc	BUG/MINOR: sched: properly account for the CPU time of dying tasks When task profiling is enabled, the scheduler can measure and report the cumulated time spent in each task and their respective latencies. But this was wrong for tasks with few wakeups as well as for self-waking ones, because the call date needed to measure how long it takes to process the task is retrieved in the task itself (->wake_date was turned to the call date), and we could face two conditions: - a new wakeup while the task is executing would reset the ->wake_date field before returning and make abnormally low values being reported; that was likely the case for task�run_applet for self-waking applets; - when the task dies, NULL is returned and the call date couldn't be retrieved, so that CPU time was not being accounted for. This was particularly visible with process_stream() which is usually called only twice per request, and whose time was systematically halved. The cleanest solution here is to keep in mind that the scheduler already uses quite a bit of local context in th_ctx, and place the intermediary values there so that they cannot vanish. The wake_date has to be reset immediately once read, and only its copy is used along the function. Note that this must be done both for tasks and tasklet, and that until recently tasklets were also able to report wrong values due to their sole dependency on TH_FL_TASK_PROFILING between tests. One nice benefit for future improvements is that such information will now be available from the task without having to be stored into the task itself anymore. Since the tasklet part was computed on wrapping 32-bit arithmetics and the task one was on 64-bit, the values were now consistently moved to 32-bit as it's already largely sufficient (4s spent in a task is more than twice what the watchdog would tolerate). Some further cleanups might be necessary, but the patch aimed at staying minimal. Task profiling output after 1 million HTTP request previously looked like this: Tasks activity: function calls cpu_tot cpu_avg lat_tot lat_avg h1_io_cb 2012338 4.850s 2.410us 12.91s 6.417us process_stream 2000136 9.594s 4.796us 34.26s 17.13us sc_conn_io_cb 2000135 1.973s 986.0ns 30.24s 15.12us h1_timeout_task 137 - - 2.649ms 19.34us accept_queue_process 49 152.3us 3.107us 321.7yr 6.564yr main+0x146430 7 5.250us 750.0ns 25.92us 3.702us srv_cleanup_idle_conns 1 559.0ns 559.0ns 918.0ns 918.0ns task_run_applet 1 - - 2.162us 2.162us Now it looks like this: Tasks activity: function calls cpu_tot cpu_avg lat_tot lat_avg h1_io_cb 2014194 4.794s 2.380us 13.75s 6.826us process_stream 2000151 20.01s 10.00us 36.04s 18.02us sc_conn_io_cb 2000148 2.167s 1.083us 32.27s 16.13us h1_timeout_task 198 54.24us 273.0ns 3.487ms 17.61us accept_queue_process 52 158.3us 3.044us 409.9us 7.882us main+0x1466e0 18 16.77us 931.0ns 63.98us 3.554us srv_cleanup_toremove_conns 8 282.1us 35.26us 546.8us 68.35us srv_cleanup_idle_conns 3 149.2us 49.73us 8.131us 2.710us task_run_applet 3 268.1us 89.38us 11.61us 3.871us Note the two-fold difference on process_stream(). This feature is essentially used for debugging so it has extremely limited impact. However it's used quite a bit more in bug reports and it would be desirable that at least 2.6 gets this fix backported. It depends on at least these two previous patches which will then also have to be backported: MINOR: task: permanently enable latency measurement on tasklets CLEANUP: task: rename ->call_date to ->wake_date	2022-09-08 14:19:15 +02:00
Willy Tarreau	04e50b3d32	CLEANUP: task: rename ->call_date to ->wake_date This field is misnamed because its real and important content is the date the task was woken up, not the date it was called. It temporarily holds the call date during execution but this remains confusing. In fact before the latency measurements were possible it was indeed a call date. Thus is will now be called wake_date. This change is necessary because a subsequent fix will require the introduction of the real call date in the thread ctx.	2022-09-08 14:19:15 +02:00
Willy Tarreau	768c2c5678	MINOR: task: permanently enable latency measurement on tasklets When tasklet latency measurement was enabled in 2.4 with commit `b2285de04` ("MINOR: tasks: also compute the tasklet latency when DEBUG_TASK is set"), the feature was conditionned on DEBUG_TASK because the field would add 8 bytes to the struct tasklet. This approach was not a very good idea because the struct ends on an int anyway thus it does finish with a 32-bit hole regardless of the presence of this field. What is true however is that adding it turned a 64-byte struct to 72-byte when caller debugging is enabled. This patch revisits this with a minor change. Now only the lowest 32 bits of the call date are stored, so they always fit in the remaining hole, and this allows to remove the dependency on DEBUG_TASK. With debugging off, we're now seeing a 48-byte struct, and with debugging on it's exactly 64 bytes, thus still exactly one cache line. 32 bits allow a latency of 4 seconds on a tasklet, which already indicates a completely dead process, so there's no point storing the upper bits at all. And even in the event it would happen once in a while, the lost upper bits do not really add any value to the debug reports. Also, now one tasklet wakeup every 4 billion will not be sampled due to the test on the value itself. Similarly we just don't care, it's statistics and the measurements are not 9-digit accurate anyway.	2022-09-08 14:19:15 +02:00
Frédéric Lécaille	614742b79c	MINOR: quic: No TRACE_LEAVE() in retrieve_qc_conn_from_cid() This macro was confused with TRACE_ENTER(). Should be backported to 2.6.	2022-09-07 15:59:43 +02:00
Frédéric Lécaille	449804e27d	MINOR: quic: Add traces about sent or resent TX frames Very useful to help in debugging issues, especially during retransmissions. Should be backported to 2.6	2022-09-07 15:59:29 +02:00
William Lallemand	70a6e637b4	MINOR: quic: add QUIC support when no client_hello_cb Add QUIC support to the ssl_sock_switchctx_cbk() variant used only when no client_hello_cb is available. This could be used with libreSSL implementation of QUIC for example. It also works with quictls when HAVE_SSL_CLIENT_HELLO_CB is removed from openss-compat.h	2022-09-07 11:33:28 +02:00
William Lallemand	373ce73695	BUILD: quic: fix the #ifdef in ssl_quic_initial_ctx() As done on with ssl_sock_initial_ctx(), cleanup the ifdef for the client_hello_cb and the no anti replay.	2022-09-07 11:11:59 +02:00
William Lallemand	4b7938d160	BUILD: ssl: fix the ifdef mess in ssl_sock_initial_ctx ssl_sock_initial_ctx uses the wrong #ifdef to check the availability of the client_hello_cb. Cleanup the #ifdef, add comments and indentation.	2022-09-07 10:54:17 +02:00
William Lallemand	e6ec626ac5	BUILD: quic: enable early data only with >= openssl 1.1.1 Disable the early data in the QUIC code when not built with openssl >= 1.1.1. LibreSSL 3.6.0 is impacted.	2022-09-07 09:33:46 +02:00
William Lallemand	844009d77a	BUILD: ssl: fix ssl_sock_switchtx_cbk when no client_hello_cb When building HAProxy with USE_QUIC and libressl 3.6.0, the ssl_sock_switchtx_cbk symbol is not found because libressl does not implement the client_hello_cb. A ssl_sock_switchtx_cbk version for the servername callback is available but wasn't exported correctly.	2022-09-07 09:33:46 +02:00
Fr�d�ric L�caille	2be0ac55e1	BUG/MINOR: quic: Possible crash when verifying certificates This verification is done by ssl_sock_bind_verifycbk() which is set at different locations in the ssl_sock.c code . About QUIC connections, there are a lot of chances the connection object is not initialized when entering this function. What must be accessed is the SSL object to retrieve the connection or quic_conn objects, then the bind_conf object of the listener. If the connection object is not found, we try to find the quic_conn object. Modify ssl_sock_dump_errors() interface which takes a connection object as parameter to also passed a quic_conn object as parameter. Again this function try first to access the connection object if not NULL or the quic_conn object if not. There is a remaining thing to do for QUIC: store the certificate verification error code as it is currently stored in the connection object. This error code is at least used by the "bc_err" and "fc_err" sample fetches. There are chances this bug is in relation with GH #1851. Thank you to @tasavis for the report. Must be merged into 2.6.	2022-09-06 20:42:02 +02:00
Christopher Faulet	a9e934bbd1	BUG/MINOR: h1: Support headers case adjustment for TCP proxies On frontend side, "h1-case-adjust-bogus-client" option is now supported in TCP mode. It is important to be able to adjust the case of response headers when a connection is routed to an HTTP backend. In this case, the client connection is upgraded to H1. On backend side, "h1-case-adjust-bogus-server" option is now also supported in TCP mode to be able to perform HTTP health-checks with a case adjustment of the request headers. This patch should be backported as far as 2.0.	2022-09-06 18:23:14 +02:00
Christopher Faulet	4b5f3029bc	MINOR: http-check: Remove support for headers/body in "option httpchk" version This trick is deprecated since the health-check refactoring, It is now invalid. It means the following line will trigger an error during the configuration parsing: option httpchk OPTIONS * HTTP/1.1\r\nHost:\ www It must be replaced by: option httpchk OPTIONS * HTTP/1.1 http-check send hdr Host www	2022-09-06 18:23:14 +02:00
Fr�d�ric L�caille	6aec1f380e	BUG/MINOR: quic: Possible crash with "tls-ticket-keys" on QUIC bind lines ssl_tlsext_ticket_key_cb() is called when "tls-ticket-keys" option is used on a "bind" line. It needs to have an access to the TLS ticket keys which have been stored into the listener bind_conf struct. The fix consists in nitializing the <ref> variable (references to TLS secret keys) the correct way when this callback is called for a QUIC connection. The bind_conf struct is store into the quic_conn object (QUIC connection). This issue may be in relation with GH #1851. Thank you for @tasavis for the report. Must be backported to 2.6.	2022-09-06 17:56:53 +02:00
Fr�d�ric L�caille	025945f12c	BUG/MINOR: quic: Retransmitted frames marked as acknowledged Obviously, frames which are duplicated from others must not be retransmitted if the original frame they were duplicated from was already acknowledged. This should have been detected by qc_build_frms() which skips such frames, except if the QUIC xprt does really bad things which are not supported by the upper layer. This will have to be checked with Amaury. To prevent the retransmision of these frames which leads to crashes as reported by hpn0t0ad this gdb backtrace in GH #1809 where the frame builder tries to copy a huge number of bytes to the packet buffer: Thread 7 (Thread 0x7fddf373a700 (LWP 13)): #0 __memmove_sse2_unaligned_erms () at ../sysdeps/x86_64/multiarch/memmove-vec-unaligned-erms.S:520 No locals. #1 0x000055b17435705e in quic_build_stream_frame (buf=0x7fddf372ef78, end=<optimized out>, frm=0x7fdde08d3470, conn=<optimized out>) at src/quic_frame.c:515 to_copy = 18446697703428890384 stream = 0x7fdde08d3490 wrap = <optimized out> which matches this part of quic_frame.c code: wrap = (const unsigned char )b_wrap(stream->buf); if (stream->data + stream->len > wrap) { size_t to_copy = wrap - stream->data; memcpy(buf, stream->data, to_copy); *buf += to_copy; we release as soon as possible the impacted frames as there is really no need to retransmit such frames. Thank you to @hpn0t0ad for having provided us with useful traces in github issue #1809. Must be backported in 2.6.	2022-09-06 14:23:52 +02:00
Brad Smith	ef9d594839	MINOR: Revert part of clarifying samples support per os commit Commit `5c83e3a156` made some adjustments to clarify which TCP_INFO information is supported by each respective OS. There was a comment like so.. Note that fc_rtt and fc_rttvar are supported on any OS that has TCP_INFO, not just linux/freebsd/netbsd, so we continue to expose them unconditionally. But the diff didn't do so in a consistent manner.	2022-09-03 06:11:08 +02:00
Willy Tarreau	6a03a0d86d	BUG/MINOR: http-act: initialize http fmt head earlier In github issue #1850, Christian Ruppert reported a case of crash in 2.6 when failing to parse some http rules. This started to happen with 2.6 commit `dd7e6c6` ("BUG/MINOR: http-rules: completely free incorrect TCP rules on error") but has some of its roots in 2.2 commit `2eb539687` ("MINOR: http-rules: Add release functions for existing HTTP actions"). The cause is that when the release function is set for HTTP actions, the rule->arg.http.fmt list head is not yet initialized, hence is NULL, thus the release function crashes when it tries to iterate over it. In fact this code was initially not written with the perspective of releasing such elements upon error, so the arg list initialization happened after error checking. This patch just moves the list initialization just after setting the release pointer and that's OK. This patch must be backported to 2.6 since the problem is visible there. It could be backported to 2.5 but the issue is not triggered there without the first mentioned patch above that landed in 2.6, so it will not bring any obvious benefit.	2022-09-02 19:24:12 +02:00
Willy Tarreau	e6f389d1a5	MINOR: mux-h1: provide a "show_sd" helper to output stream debugging info With this, it now becomes possible to see the state of each H1 stream from "show sess all". Example (added lines highlighted with '>'): 0x7fc9b40460a0: [02/Sep/2022:16:29:31.267228] id=49 proto=tcpv4 source=127.0.0.1:53548 flags=0xc4a, conn_retries=0, conn_exp=<NEVER> conn_et=0x000 srv_conn=0x2dc4b20, pend_pos=(nil) waiting=0 epoch=0 frontend=decrypt (id=2 mode=http), listener=? (id=3) addr=127.0.0.1:8001 backend=decrypt (id=2 mode=http) addr=127.0.0.1:25168 server=httpterm (id=1) addr=127.0.0.1:8000 task=0x7fc9b4046490 (state=0x00 nice=0 calls=4 rate=0 exp=3s tid=7(1/7) age=2s) txn=0x7fc9b4046650 flags=0x3000 meth=1 status=200 req.st=MSG_DONE rsp.st=MSG_DATA req.f=0x4c rsp.f=0x0d scf=0x7fc9b4046030 flags=0x00000080 state=EST endp=CONN,0x7fc9b4041f00,0x02804001 sub=1 > h1s=0x7fc9b4041f00 h1s.flg=0x104010 .sd.flg=0x2804001 .req.state=MSG_DONE .res.state=MSG_DATA > .meth=GET status=200 .sd.flg=0x02804001 .sc.flg=0x00000080 .sc.app=0x7fc9b40460a0 > .subs=0x7fc9b4046040(ev=1 tl=0x7fc9b4046540 tl.calls=9 tl.ctx=0x7fc9b4046030 tl.fct=sc_conn_io_cb) > h1c=0x7fc9b402b3f0 h1c.flg=0x302200 .sub=1 .ibuf=0@(nil)+0/0 .obuf=0@(nil)+0/0 co0=0x7fc9bc02e740 ctrl=tcpv4 xprt=RAW mux=H1 data=STRM target=LISTENER:0x2dc3c40 flags=0x00000300 fd=79 fd.state=421 updt=0 fd.tmask=0x80 scb=0x7fc9b4046590 flags=0x00000011 state=EST endp=CONN,0x7fc9b4048660,0x02840001 sub=0 > h1s=0x7fc9b4048660 h1s.flg=0x4010 .sd.flg=0x2840001 .req.state=MSG_DONE .res.state=MSG_DATA > .meth=GET status=200 .sd.flg=0x02840001 .sc.flg=0x00000011 .sc.app=0x7fc9b40460a0 .subs=(nil) > h1c=0x7fc9b4048490 h1c.flg=0x80002200 .sub=0 .ibuf=0@(nil)+0/0 .obuf=0@(nil)+0/0 co1=0x7fc9b4048270 ctrl=tcpv4 xprt=RAW mux=H1 data=STRM target=SERVER:0x2dc4b20 flags=0x00000300 fd=131 fd.state=10122 updt=0 fd.tmask=0x80 req=0x7fc9b40460c0 (f=0x49c40080 an=0x8000 pipe=0 tofwd=0 total=56) an_exp=<NEVER> rex=<NEVER> wex=<NEVER> buf=0x7fc9b40460c8 data=(nil) o=0 p=0 i=0 size=0 htx=0xdd90a0 flags=0x0 size=0 data=0 used=0 wrap=NO extra=0 res=0x7fc9b4046120 (f=0x80070202 an=0x4000000 pipe=0 tofwd=-1 total=603840788) an_exp=<NEVER> rex=<NEVER> wex=<NEVER> buf=0x7fc9b4046128 data=(nil) o=0 p=0 i=0 size=0 htx=0xdd90a0 flags=0x0 size=0 data=0 used=0 wrap=NO extra=0	2022-09-02 16:43:25 +02:00
Willy Tarreau	7079c0fbbd	MINOR: mux-h1: split "show_fd" into connection and stream We now have two functions, one for dumping connections and the other one for dumping the streams. This will permit to use it from show_sd. A few optional line breaks were inserted where relevant to keep lines homogenous when a prefix is passed.	2022-09-02 16:43:25 +02:00
Willy Tarreau	b4a4feee87	MINOR: mux-quic: provide a "show_sd" helper to output stream debugging info It's very limited but at least provides the very basic info about QCS and QCC when issuing "show sess all": scf=0x7fa9642394a0 flags=0x00000080 state=EST endp=CONN,0x7fa9642351f0,0x02001001 sub=3 > qcs=0x7fa9642351f0 .flg=0x5 .id=396 .st=HCR .ctx=0x7fa9642353f0, .err=0 > qcc=0x7fa96405ce20 .flg=0 .nbsc=100 .nbhreq=100, .task=0x7fa964054260 co0=0x7fa96405cd50 ctrl=quic4 xprt=QUIC mux=QUIC data=STRM target=LISTENER:0x328c530 flags=0x00200300 fd=-1 fd.state=00 updt=0 fd.tmask=0x0 It will need to be improved but it's better than nothing already. This should be backported to 2.6 if the other dumps are backported.	2022-09-02 16:43:25 +02:00
Willy Tarreau	7051f73efe	MINOR: mux-h2: insert line breaks in "show sess all" output for legibility h2s and h2c were extremely long in the "show sess all" output, around 300 chars each. This adds a few line breaks to improve legibility, there are now 3 lines for each, which are around the same length as the other ones while keeping a natural arrangement. E.g (lines highlighted with '>'): 0x7faad8144f80: [02/Sep/2022:15:49:40.171620] id=105283 proto=tcpv4 source=127.0.0.1:42942 flags=0x100c4a, conn_retries=0, conn_exp=<NEVER> conn_et=0x000 srv_conn=0x1f44b20, pend_pos=(nil) waiting=0 epoch=0 frontend=decrypt (id=2 mode=http), listener=? (id=3) addr=127.0.0.1:8001 backend=decrypt (id=2 mode=http) addr=127.0.0.1:18144 server=httpterm (id=1) addr=127.0.0.1:8000 task=0x7faad812b7c0 (state=0x00 nice=0 calls=2 rate=0 exp=4s tid=7(1/7) age=0s) txn=0x7faad81453e0 flags=0x43000 meth=1 status=200 req.st=MSG_DONE rsp.st=MSG_DATA req.f=0x4c rsp.f=0x0d scf=0x7faad81625d0 flags=0x00000080 state=EST endp=CONN,0x7faad811d380,0x02805001 sub=1 > h2s=0x7faad811d380 h2s.id=2113 .st=HCR .flg=0x207001 .rxbuf=0@(nil)+0/0 > .sc=0x7faad81625d0(.flg=0x00000080 .app=0x7faad8144f80) .sd=0x7faad8119dc0(.flg=0x02805001) > .subs=0x7faad81625e0(ev=1 tl=0x7faad86d6500 tl.calls=4 tl.ctx=0x7faad81625d0 tl.fct=sc_conn_io_cb) > h2c=0x7faad802c640 h2c.st0=FRH .err=0 .maxid=2157 .lastid=-1 .flg=0x0600 .nbst=1 .nbsc=1 > .fctl_cnt=0 .send_cnt=0 .tree_cnt=1 .orph_cnt=0 .sub=1 .dsi=2157 .dbuf=0@(nil)+0/0 > .msi=-1 .mbuf=[6..6\|32],h=[0@(nil)+0/0],t=[0@(nil)+0/0] co0=0x7faae402efc0 ctrl=tcpv4 xprt=RAW mux=H2 data=STRM target=LISTENER:0x1f43c40 flags=0x00000300 fd=95 fd.state=121 updt=0 fd.tmask=0x80 scb=0x7faad8145370 flags=0x00000011 state=EST endp=CONN,0x7faad8115630,0x02840001 sub=1 co1=0x7faad86c0730 ctrl=tcpv4 xprt=RAW mux=H1 data=STRM target=SERVER:0x1f44b20 flags=0x00000300 fd=1656 fd.state=10121 updt=0 fd.tmask=0x80 req=0x7faad8144fa0 (f=0x49c40000 an=0x8000 pipe=0 tofwd=0 total=110) an_exp=<NEVER> rex=<NEVER> wex=<NEVER> buf=0x7faad8144fa8 data=(nil) o=0 p=0 i=0 size=0 htx=0xdd90a0 flags=0x0 size=0 data=0 used=0 wrap=NO extra=0 res=0x7faad8145000 (f=0x80040202 an=0x4000000 pipe=0 tofwd=-1 total=60365) an_exp=<NEVER> rex=<NEVER> wex=<NEVER> buf=0x7faad8145008 data=(nil) o=0 p=0 i=0 size=0 htx=0xdd90a0 flags=0x0 size=0 data=0 used=0 wrap=NO extra=0	2022-09-02 16:43:03 +02:00
Willy Tarreau	bf4ec6f4a0	MINOR: mux-h2: provide a "show_sd" helper to output stream debugging info With this, it now becomes possible to see the state of each H2 stream from "show sess all". Lines are still too long and need to be split, but that's for another patch.	2022-09-02 15:48:50 +02:00
Willy Tarreau	ce57777660	MINOR: muxes: add a "show_sd" helper to complete "show sess" dumps This helper will be called for muxes that provide it and will be used to let the mux provide extra information about the stream attached to a stream descriptor. A line prefix is passed in argument so that the mux is free to break long lines without breaking indent. No prefix means no line breaks should be produced (e.g. for short dumps).	2022-09-02 15:48:50 +02:00
Willy Tarreau	4e97bcc76b	MINOR: mux-h2: extract the connection dump function out of h2_show_fd() The function will be reusable to dump connections, so let's extract it.	2022-09-02 15:48:10 +02:00
Willy Tarreau	90bffa2ce3	MINOR: mux-h2: extract the stream dump function out of h2_show_fd() The function will be reusable to dump streams, so let's extract it. Note that due to "last_h2s" being originally printed as a prefix for the stream dump, now the pointer is displayed by the caller instead.	2022-09-02 15:48:10 +02:00
Willy Tarreau	714900a3c9	MINOR: debug: report applet pointer and handler in crashes when known When an appctx is found looping over itself, we report a number of info but not the pointers to the definition nor the handler, which can be quite handy in some cases. Let's add them and try to decode the symbol.	2022-09-02 15:48:10 +02:00
Willy Tarreau	410546145b	BUG/MINOR: mux-fcgi: fix the "show fd" dest buffer for the subscriber Commit `1776ffb97` ("MINOR: mux-fcgi: make the "show fd" helper also decode the fstrm subscriber when known") improved the output of "show fd" for the FCGI mux, but the output is sent to the trash buffer instead of the msg argument. It turns out that this has no effect right now as the caller passes the trash but this is risky. This should be backported to 2.4.	2022-09-02 14:23:56 +02:00
Willy Tarreau	9b6a187e26	BUG/MINOR: mux-h1: fix the "show fd" dest buffer for the subscriber Commit `150c4f8b7` ("MINOR: mux-h1: make the "show fd" helper also decode the h1s subscriber when known") improved the output of "show fd" for the H1 mux, but the output is sent to the trash buffer instead of the msg argument. It turns out that this has no effect right now as the caller passes the trash but this is risky. This should be backported to 2.4.	2022-09-02 14:23:56 +02:00
Willy Tarreau	ba7657ca0f	BUG/MINOR: mux-h2: fix the "show fd" dest buffer for the subscriber Commit `98e40b981` ("MINOR: mux-h2: make the "show fd" helper also decode the h2s subscriber when known") improved the output of "show fd" for the H2 mux, but the output is sent to the trash buffer instead of the msg argument. It turns out that this has no effect right now as the caller passes the trash but this is risky. This should be backported to 2.4.	2022-09-02 14:23:56 +02:00
Willy Tarreau	df3231c74a	MEDIUM: httpclient: enable ALPN support on outgoing https connections Since everything is available for this, let's enable ALPN with the usual "h2,http/1.1" on the https server. This will allow HTTPS requests to use HTTP/2 when available. It may be needed to permit to disable this (or to set the string) in case some client code explicitly checks for the "HTTP/1.1" string, but since httpclient is quite young it's unlikely that such code already exists.	2022-09-02 13:54:30 +02:00
Willy Tarreau	f80713ba8e	BUG/MINOR: httpclient: keep-alive was accidentely disabled The servers were not set with default settings, meaning that a few settings including the pool_max_delay were not set, thus disabling connection pools, which is the cause of the fact that keep-alive was disabled as reported in issue #1831. There might possibly be other issues pending since all these fields were left to zero. Note that this patch alone will not fix keep-alive because the applet does not enforce SE_FL_NOT_FIRST and relies on the default http-reuse safe, thus if servers are not shared, all requests are considered first ones and do not reuse existing connections. In 2.7, commit `ecb40b2c3` ("MINOR: backend: always satisfy the first req reuse rule with l7 retries") addressed this in a more elegant way by fixing http-reuse to take into account the fact that properly configured l7 retries provide exactly the capability that reuse safe was trying to cover, and this patch is suitable for backporting. This patch should be backported to 2.6 only.	2022-09-02 11:48:01 +02:00
Willy Tarreau	6486ff8cab	BUG/MINOR: httpclient: only ask for more room on failed writes There's a tiny issue in the I/O handler by which both a failed request emission and missing response data will want to subscribe for more room on output. That's not correct in that only the case where the request buffer is full should cause this, the other one should just wait for incoming data. This could theoretically cause spurious wakeups at certain key points (e.g. connect() time maybe) though this could not be reproduced but better fix this while it's easy enough. It doesn't seem necessary to backport it right now, though this may have to in case a concrete reproducible case is discovered.	2022-09-02 11:42:50 +02:00
Willy Tarreau	b48292068b	BUG/MEDIUM: httpclient: always detach the caller before self-killing If the caller dies before the server responds, the httpclient can crash in hc_cli_res_end_cb() when unregistering because it dereferences hc->caller which was already freed during the caller's unregistration. The easiest way to reproduce it is by sending twice the following request on the same CLI connection in expert mode, with httpterm running on local port 8000: httpclient GET http://127.0.0.1:8000/?t=600 Note the 600ms delay that's larger than socat's default 500. The code checks for a NULL everywhere hc->caller is used, but the NULL was forgotten in this specific case. It must be placed in the second half of httpclient_stop_and_destroy() which is responsible for signaling the client that the caller leaves. This must be backported to 2.6.	2022-09-02 11:19:07 +02:00
Willy Tarreau	d8a44d0b24	BUG/MINOR: h2: properly set the direction flag on HTX response In 1.9-dev, a new flag was introduced on the start line with commit `f1ba18d7b` ("MEDIUM: htx: Don't rely on h1_sl anymore except during H1 header parsing") to designate a response message: HTX_SL_F_IS_RESP. Unfortunately as it was done in parallel to the mux_h2 support for the backend, it was never integrated there. It was not used by then so this remained unnoticed for a while. However the http_client now uses it, and missing that flag prevents it from using the H2 mux, so let's properly add it. There's no point in backporting this far away, but since the http_client is fully operational in 2.6 it would make sense to backport this fix at least there to secure the code.	2022-09-02 11:19:07 +02:00
Fr�d�ric L�caille	a1075209c7	BUG/MINOR: quic: Frames leak during retransmissions The frame which are retransmitted by qc_dgrams_retransmit() are duplicated from sent but not acknowledged packets and added to local frames lists. Some may not have been sent. If not replaced somewhere (linked to the connection) they are lost for ever (leak). We splice the list remaining contents to the packets number space frame list to avoid such a situation. Must be backported to 2.6.	2022-09-02 08:47:38 +02:00
Fr�d�ric L�caille	a777ee36f6	MINOR: quic: Trace typo fix in qc_release_frm() Grammar fix without any impact.	2022-09-02 08:47:38 +02:00
Fr�d�ric L�caille	26236f5a5d	MINOR: quic: Add TX frames addresses to traces to several trace events This should be useful to diagnose TX frames related issues.	2022-09-02 08:47:38 +02:00
Fr�d�ric L�caille	b866c69f4f	BUG/MINOR: quic: Do not ack when probing <force_ack> boolean variable passed to qc_do_build_pkt() which builds a clear packet is there to force this function to build an ACK frame regardless of others conditions. This is used during handshake, when we acknowledge every handshake packets received. This variable was already taken into an account by the local variable <must_ack> which is there at least to ignore any other conditions than this one: "are we building a probing packet?". Indeed we do not want to add ACK frames when we probe the peers. This is to have more chances to embed the new duplicated frames into another packets without splitting them. So, the test on <force_ack> boolean value is useless, silly and brakes the rule which consists in not acknowledging when probing. Must be backported to 2.6.	2022-09-02 08:47:38 +02:00
Willy Tarreau	ecb40b2c38	MINOR: backend: always satisfy the first req reuse rule with l7 retries The "first req" rule consists in not delivering a connection's first request to a connection that's not known for being safe so that we don't deliver a broken page to a client if the server didn't intend to keep it alive. That's what's used by "http-reuse safe" particularly. But the reason this rule was created was precisely because haproxy was not able to re-emit the request to the server in case of connection breakage, which is precisely what l7 retries later brought. As such, there's no reason for enforcing this rule when l7 retries are properly enabled because such a blank page will trigger a retry and will not be delivered to the client. This patch simply checks that the l7 retries are enabled for the 3 cases that can be triggered on a dead or dying connection (failure, empty, and timeout), and if all 3 are enabled, then regular idle connections can be reused. This could almost be marked as a bug fix because a lot of users relying on l7 retries do not necessarily think about using http-reuse always due to the recommendation against it in the doc, while the protection that the safe mode offers is never used in that mode, and it forces the http client not to reuse existing persistent connections since it never sets the "not first" flag. It could also be decided that the protection is not used either when the origin is an applet, as in this case this is internal code that we can decide to let handle the retry by itself (all info are still present). But at least the httpclient will be happy with this alone. It would make sense to backport this at least to 2.6 in order to let the httpclient reuse connections, maybe to older releases if some users report low reuse counts.	2022-09-01 20:52:29 +02:00
Willy Tarreau	4d1ff11f05	BUG/MEDIUM: mux-h1: always use RST to kill idle connections in pools When idle H1 connections cannot be stored into a server pool or are later evicted, they're often seen closed with a FIN then an RST. The problem is that this is sufficient to leave them in TIME_WAIT in the local sockets table and port exhaustion may happen. The reason is that in h1_release() we rely on h1_shutw_conn() which itself decides whether to close in silent or normal mode only based on the H1C_F_ST_SILENT_SHUT flag. This flag is only set by h1_shutw() based on the requested mode. But when the connection is in the idle list, the mode ought to always be silent. What this patch does is to set the flag before trying to add to the idle list, and remove it after removing from the idle list. This way if the connection fails to be added or has to be killed, it's closed with an RST. This must be backported as far as 2.4. It's not sure whether older versions need an equivalent.	2022-09-01 20:52:29 +02:00
Christopher Faulet	f348ecd67a	BUG/MINOR: regex: Properly handle PCRE2 lib compiled without JIT support The PCRE2 JIT support is buggy. If HAProxy is compiled with USE_PCRE2_JIT option while the PCRE2 library is compiled without the JIT support, any matching will fail because pcre2_jit_compile() return value is not properly handled. We must fall back on pcre2_match() if PCRE2_ERROR_JIT_BADOPTION error is returned. This patch should fix the issue #1848. It must be backported as far as 2.4.	2022-09-01 19:34:46 +02:00
Willy Tarreau	32872db605	MINOR: sink/ring: rotate non-empty file-backed contents only If the service is rechecked before a reload, that may cause the config to be parsed twice and file-backed rings to be lost. Here we make sure that such a ring does contain information before deciding to rotate it. This way the first process starting after some writes will cause a rotate but not subsequent ones until new writes are applied. An attempt was also made to disable rotations on checks but this was a bad idea, as the ring is still initialized and this causes the contents to be lost. The choice of initializing the ring during parsing is questionable but the config check ought to be as close as possible to a real start, and we could imagine that the ring is used by some code during startup (e.g. lua). So this approach was abandonned and config checks also cause a rotation, as the purpose of this rotation is to preserve latest information against accidental removal.	2022-09-01 08:25:34 +02:00
William Lallemand	e0fa91ffe1	BUG/MINOR: ssl: leak of ckch_inst_link in ckch_inst_free() v2 ckch_inst_free() unlink the ckch_inst_link structure but never free it. It can't be fixed simply because cli_io_handler_commit_cafile_crlfile() is using a cafile_entry list to iterate a list of ckch_inst entries to free. So both cli_io_handler_commit_cafile_crlfile() and ckch_inst_free() would modify the list at the same time. In order to let the caller manipulate the ckch_inst_link, ckch_inst_free() now checks if the element is still attached before trying to detach and free it. For this trick to work, the caller need to do a LIST_DEL_INIT() during the iteration over the ckch_inst_link. list_for_each_entry was also replace by a while (!LIST_ISEMPTY()) on the head list in cli_io_handler_commit_cafile_crlfile() so the iteration works correctly, because it could have been stuck on the first detached element. list_for_each_entry_safe() is not enough to fix the issue since multiple element could have been removed. Must be backported as far as 2.5.	2022-08-31 15:24:01 +02:00
Fr�d�ric L�caille	bccbad2654	BUG/MINOR: quic: TX frames memleak Missing call to pool_free() for quic_frame objects Must be backported to 2.6.	2022-08-31 15:20:29 +02:00
Fr�d�ric L�caille	3a9b944955	MINOR: quic: Move traces about RX/TX bytes from QUIC_EV_CONN_PRSAFRM event Move these traces to QUIC_EV_CONN_SPPKTS trace event. They were displayed at a useless location. Make them displayed just after having sent a packet and when checking the anti-amplication limit. Useful to diagnose issues in relation with the recovery.	2022-08-31 15:20:24 +02:00
William Lallemand	0bfa3e7ff2	BUG/MINOR: ssl: revert two wrong fixes with ckhi_link This reverts commit `056ad01d55`. This reverts commit `ddd480cbdc`. The architecture is ambiguous here: ckch_inst_free() is detaching and freeing the "ckch_inst_link" linked list which must be free'd only from the cafile_entry side. The problem was also hidden by the fix `ddd480c` ("BUG/MEDIUM: ssl: Fix a UAF when old ckch instances are released") which change the ckchi_link inner loop by a safe one. However this can't fix entirely the problem since both __ckch_inst_free_locked() could remove several nodes in the ckchi_link linked list. This revert is voluntary reintroducing a memory leak before really fixing the problem. Must be backported in 2.5 + 2.6.	2022-08-30 18:12:28 +02:00
Christopher Faulet	ddd480cbdc	BUG/MEDIUM: ssl: Fix a UAF when old ckch instances are released When old chck instances is released at the end of "commit ssl ca-file" or "commit ssl crl-file" commands, the link is released. But we walk through the list using the unsafe macro. list_for_each_entry_safe() must be used. This bug was introduced by commit `056ad01d5` ("BUG/MINOR: ssl: leak of ckch_inst_link in ckch_inst_free()"). Thus this patch must be backported as far as 2.5.	2022-08-30 16:27:51 +02:00
Christopher Faulet	f611248d8c	BUG/MINOR: tcpcheck: Disable QUICKACK for default tcp-check (with no rule) The commit `871dd8211` ("BUG/MINOR: tcpcheck: Disable QUICKACK only if data should be sent after connect") introduced a regression. It removes the test on the next rule to be able to disable TCP_QUICKACK when only a connect is performed (so no next rule). This patch must be backported as far as 2.2.	2022-08-30 10:31:16 +02:00
William Lallemand	056ad01d55	BUG/MINOR: ssl: leak of ckch_inst_link in ckch_inst_free() ckch_inst_free() unlink the ckch_inst_link structure but never free it. It can cause a memory leak upon a ckch_inst_free() done with CLI operation. Bug introduced by commit `4458b97` ("MEDIUM: ssl: Chain ckch instances in ca-file entries"). Must be backported as far as 2.5.	2022-08-29 18:53:34 +02:00
William Lallemand	946580e17a	BUG/MINOR: ssl: fix deinit of the ca-file tree Commit `b0c4827` ("BUG/MINOR: ssl: free the cafile entries on deinit") introduced a double free. The node was never removed from the tree before its free. Fix issue #1836. Must be backported where `b0c4827` was backported. (2.6 for now).	2022-08-29 18:51:39 +02:00
Fr�d�ric L�caille	3a56137048	MINOR: quic: Add a trace to distinguish the datagram from the packets inside Without such a trace, we do not know when a datagram is sent. Only trace for the packets inside the datagrams were displayed. Must be backported to 2.6.	2022-08-29 18:46:40 +02:00
Fr�d�ric L�caille	c242832af3	BUG/MINOR: quic: Missing header protection AES cipher context initialisations (draft-v2) This bug arrived with this commit: "MINOR: quic: Add reusable cipher contexts for header protection" haproxy could crash because of missing cipher contexts initializations for the header protection and draft-v2 Initial secrets. This was due to the fact that these initialization both for RX and TX secrets were done outside of qc_new_isecs(). The role of this function is definitively to initialize these cipher contexts in addition to the derived secrets. Indeed this function is called by qc_new_conn() which initializes the connection but also by qc_conn_finalize() which also calls qc_new_isecs() in case of a different QUIC version was negotiated by the peers from the one used by the client for its first Initial packet. This was reported by "v2" QUIC interop test with at least picoquic as client. Must be backported to 2.6.	2022-08-29 18:46:40 +02:00
Willy Tarreau	c6fc77404e	MINOR: raw-sock: don't try to send if an error was already reported There's no point trying to send() on a socket on which an error was already reported. This wastes syscalls. Till now it was possible to occasionally see an attempt to sendto() after epoll_wait() had reported EPOLLERR.	2022-08-29 18:45:27 +02:00
Willy Tarreau	2c30de3b90	BUG/MINOR: epoll: do not actively poll for Rx after an error In 2.2, commit `5d7dcc2a8` ("OPTIM: epoll: always poll for recv if neither active nor ready") was added to compensate for the fact that our iocbs are almost always asynchronous now and do not have the opportunity to update the FD correctly. As such, they just perform a wakeup, the FD is turned to inactive, the tasklet wakes up, performs the I/O, updates the FD, most of the time this is done withing the same polling loop, and the update cancels itself in the poller without having to switch the FD off then on. The issue was that when deciding to claim an FD was active for reads if it was active for writes, we forgot one situation that unfortunately causes excessive wakeups: dealing with errors. Indeed, errors are reported and keep ringing as long as the FD is active for sending even if the consumer disabled the FD for receiving. Usually this only causes one extra wakeup for the time it takes to consider a potential write subscriber and to call it, though with many tasks in a run queue, it can last a bit longer and be reported more often. The fix consists in checking that we really want to get more receive events on this FD, that is: - that no prevous EPOLLERR was reported - that the FD doesn't carry a sticky error - that the FD is not shut for reads With this, after the last epoll_wait() reports EPOLLERR, one last recv() is performed to flush pending data and the FD is immediately unregistered. It's probably not needed to backport this as its effects are not much visible, though it should not harm. Before, EPOLLERR was seen twice: accept4(4, {sa_family=AF_INET, sin_port=htons(22314), sin_addr=inet_addr("127.0.0.1")}, [128 => 16], SOCK_NONBLOCK) = 8 accept4(4, 0x261b160, [128], SOCK_NONBLOCK) = -1 EAGAIN (Resource temporarily unavailable) recvfrom(8, "POST / HTTP/1.1\r\nConnection: close\r\nTransfer-encoding: chunk"..., 16320, 0, NULL, NULL) = 66 socket(AF_INET, SOCK_STREAM, IPPROTO_IP) = 9 connect(9, {sa_family=AF_INET, sin_port=htons(8002), sin_addr=inet_addr("127.0.0.1")}, 16) = -1 EINPROGRESS (Operation now in progress) epoll_ctl(3, EPOLL_CTL_ADD, 8, {events=EPOLLIN\|EPOLLRDHUP, data={u32=8, u64=8}}) = 0 epoll_ctl(3, EPOLL_CTL_ADD, 9, {events=EPOLLIN\|EPOLLOUT\|EPOLLRDHUP, data={u32=9, u64=9}}) = 0 epoll_wait(3, [{events=EPOLLOUT, data={u32=9, u64=9}}], 200, 355) = 1 recvfrom(9, 0x25cfb30, 16320, 0, NULL, NULL) = -1 EAGAIN (Resource temporarily unavailable) sendto(9, "POST / HTTP/1.1\r\ntransfer-encoding: chunked\r\n\r\n", 47, MSG_DONTWAIT\|MSG_NOSIGNAL, NULL, 0) = 47 epoll_ctl(3, EPOLL_CTL_MOD, 9, {events=EPOLLIN\|EPOLLRDHUP, data={u32=9, u64=9}}) = 0 epoll_wait(3, [{events=EPOLLIN\|EPOLLERR\|EPOLLHUP\|EPOLLRDHUP, data={u32=9, u64=9}}], 200, 354) = 1 recvfrom(9, "HTTP/1.1 200 OK\r\ncontent-length: 0\r\nconnection: close\r\n\r\n", 16320, 0, NULL, NULL) = 57 sendto(8, "HTTP/1.1 200 OK\r\ncontent-length: 0\r\nconnection: close\r\n\r\n", 57, MSG_DONTWAIT\|MSG_NOSIGNAL, NULL, 0) = 57 ->epoll_wait(3, [{events=EPOLLIN\|EPOLLERR\|EPOLLHUP\|EPOLLRDHUP, data={u32=9, u64=9}}], 200, 354) = 1 epoll_ctl(3, EPOLL_CTL_DEL, 9, 0x7ffe0b65fb24) = 0 epoll_wait(3, [{events=EPOLLIN, data={u32=8, u64=8}}], 200, 354) = 1 recvfrom(8, "A\n0123456789\r\n0\r\n\r\n", 16320, 0, NULL, NULL) = 19 close(9) = 0 close(8) = 0 After, EPOLLERR is only seen only once, with one less call to epoll_wait(): accept4(4, {sa_family=AF_INET, sin_port=htons(22362), sin_addr=inet_addr("127.0.0.1")}, [128 => 16], SOCK_NONBLOCK) = 8 accept4(4, 0x20d0160, [128], SOCK_NONBLOCK) = -1 EAGAIN (Resource temporarily unavailable) recvfrom(8, "POST / HTTP/1.1\r\nConnection: close\r\nTransfer-encoding: chunk"..., 16320, 0, NULL, NULL) = 66 socket(AF_INET, SOCK_STREAM, IPPROTO_IP) = 9 connect(9, {sa_family=AF_INET, sin_port=htons(8002), sin_addr=inet_addr("127.0.0.1")}, 16) = -1 EINPROGRESS (Operation now in progress) epoll_ctl(3, EPOLL_CTL_ADD, 8, {events=EPOLLIN\|EPOLLRDHUP, data={u32=8, u64=8}}) = 0 epoll_ctl(3, EPOLL_CTL_ADD, 9, {events=EPOLLIN\|EPOLLOUT\|EPOLLRDHUP, data={u32=9, u64=9}}) = 0 epoll_wait(3, [{events=EPOLLOUT, data={u32=9, u64=9}}], 200, 411) = 1 recvfrom(9, 0x2084b30, 16320, 0, NULL, NULL) = -1 EAGAIN (Resource temporarily unavailable) sendto(9, "POST / HTTP/1.1\r\ntransfer-encoding: chunked\r\n\r\n", 47, MSG_DONTWAIT\|MSG_NOSIGNAL, NULL, 0) = 47 epoll_ctl(3, EPOLL_CTL_MOD, 9, {events=EPOLLIN\|EPOLLRDHUP, data={u32=9, u64=9}}) = 0 epoll_wait(3, [{events=EPOLLIN\|EPOLLERR\|EPOLLHUP\|EPOLLRDHUP, data={u32=9, u64=9}}], 200, 411) = 1 recvfrom(9, "HTTP/1.1 200 OK\r\ncontent-length: 0\r\nconnection: close\r\n\r\n", 16320, 0, NULL, NULL) = 57 sendto(8, "HTTP/1.1 200 OK\r\ncontent-length: 0\r\nconnection: close\r\n\r\n", 57, MSG_DONTWAIT\|MSG_NOSIGNAL, NULL, 0) = 57 epoll_ctl(3, EPOLL_CTL_DEL, 9, 0x7ffc95d46f04) = 0 epoll_wait(3, [{events=EPOLLIN, data={u32=8, u64=8}}], 200, 411) = 1 recvfrom(8, "A\n0123456789\r\n0\r\n\r\n", 16320, 0, NULL, NULL) = 19 close(9) = 0 close(8) = 0	2022-08-29 18:45:27 +02:00
Willy Tarreau	cad42a78b8	BUG/MEDIUM: mux-h1: do not refrain from signaling errors after end of input In 2.6-dev4, a fix for truncated response was brought with commit `99bbdbcc2` ("BUG/MEDIUM: mux-h1: only turn CO_FL_ERROR to CS_FL_ERROR with empty ibuf"), trying to address the situation where an error is present at the connection level but some data are still pending to be read by the stream. However, this patch did not consider the case where the stream was no longer willing to read the pending data, resulting in a situation where some aborted transfers could lead to excessive CPU usage by causing constant stream wakeups for which no error was reported. This perfectly matches what was observed and reported in github issue #1842. It's not trivial to reproduce, but aborting HTTP/1 pipelining in the middle of transfers seems to give good results (using h2load and Ctrl-C in the middle). The fix was incorrct as the error should be held only if there were data that the stream was able to read. This is the approach taken by this patch, which also checks via SE_FL_EOI \| SE_FL_EOS that the stream will be able to consume the pending data. Note that the loop was provoked by the attempt by sc_conn_io_cb() itself to call sc_conn_send() which resulted in a write subscription in h1_subscribe() which immediately calls a tasklet_wakeup() since the event is ready, and that it is now stopped by the presence of SE_FL_ERROR that is checked in sc_conn_io_cb(). It seems that an extra check down the send() path to refrain from subscribing when the connection is in error could speed up error detection or at least avoid a risk of loops in this case, but this is tricky. In addition, there's already SE_FL_ERR_PENDING that seems more suitable for reporting when there are pending data, but similarly, it probably isn't checked well enough to be suitable for backports. FWIW the issue may (unreliably) be reproduced by chaining haproxy to httpterm and issuing: (printf "GET /?s=10g HTTP/1.1\r\n\r\n"; sleep 0.1; printf "\r\n") \| \ nc6 --half-close 0 8001 \| head -c1000000000 >/dev/null It's necessary to play with the size of the head command that's supposed to trigger the error at some point. A variant involving h2load in h1 mode and multiple pipelined streams, that is stopped with Ctrl-C also tends to work. As the fix above was backported as far as 2.0, it would be tempting to backport this one as far. However tests have shown that the oldest version that can trigger this issue is 2.5, maybe due to subtle differences in older ones, so it's probably not worth going further until an issue is reported. Note that in 2.5 and older, the SE_FL_* flags are applied on the conn_stream instead, as CS_FL_*. Special thanks go to Felipe W Damasio for providing lots of detailed data allowing to quickly spot the root cause of the problem.	2022-08-29 18:45:27 +02:00
Christopher Faulet	4a20972a95	BUG/MINOR: hlua: Rely on CF_EOI to detect end of message in HTTP applets applet:getline() and applet:receive() functions for HTTP applets must rely on the channel flags to detect the end of the message and not on HTX flags. It means CF_EOI must be used instead of HTX_FL_EOM. It is important because the HTX flag is transient. Because there is no flag on HTTP applets to save the info, it is not reliable. However CF_EOI once set is never removed. So it is safer to rely on it. Otherwise, the call to these functions hang. This patch must be backported as far as 2.4.	2022-08-29 15:37:17 +02:00
Christopher Faulet	b372f16d35	BUG/MEDIUM: peers: Don't start resync on reload if local peer is not up-to-date On a reload, if the previous resync was not finished, the freshly old worker must not try to start a new resync. Otherwise, it will compete with the older wokers, slowing down or blocking the resync. Only an up-to-date woker must try to perform a local resync. This patch must be backported as far as 2.0 (and maybe to 1.8 too).	2022-08-29 11:38:02 +02:00
Christopher Faulet	19a82b9495	BUG/MEDIUM: peers: Don't use resync timer when local resync is in progress When a worker is stopped, the resync timer is used to limit in time the connection stage to the new worker to perform the local resync. However, this timer must be stopped when the resync is in progress and it must be re-armed if the resync is interrupted (for instance because another reload). Otherwise, if the resync is a bit long, an old worker may be killed too early. This bug was introduce by the commit `160fff665` ("BUG/MEDIUM: peers: limit reconnect attempts of the old process on reload"). It must be backported as far as 2.0.	2022-08-29 11:38:02 +02:00
Christopher Faulet	13db4bdbc6	BUG/MEDIUM: peers: Add connect and server timeut to peers proxy Only the client timeout was set. Nothing prevent a peer applet to stall during a connect or waiting a message from a remote peer. To avoid any issue, it is important to also set connection and server timeouts. The connect timeout is set to 1s and the server timeout is set to 5s. This patch must be backported to all supported versions.	2022-08-29 11:38:02 +02:00
Christopher Faulet	42a0662910	BUG/MEDIUM: spoe: Properly update streams waiting for a ACK in async mode A bug was introduced by the commit `b042e4f6f` ("BUG/MAJOR: spoe: properly detach all agents when releasing the applet"). The fix is not correct. We really want to known if the released appctx is the last one or not. It is important when async mode is used. If there are still running applets, we just need to remove the reference on the current applet from streams in the async waiting queue. With the commit above, in async mode, if there are still running applets, it will work as expected. Otherwise a processing timeout will be reported for all these streams. So it is not too bad. But for other modes (sync and pipelining), the async waiting queue is always empty. If at least one stream is waiting to send a message, a new applet is created. It is an issue if the SPOA is unhealthy because the number of running applets may explode. However, the commit above tried to fix an issue. The bug is in fact when an new SPOE applet is created. On success, we must remove reference on the current appctx from the streams in the async waiting queue. This patch must be backported as far as 1.8.	2022-08-29 09:57:33 +02:00
Frédéric Lécaille	149c531fa1	BUG/MINOR: quic: Frames added to packets even if not built. Several frames could remain as not build into <frm_list> built by qc_build_frms() after having stopped at the first building error. So only one frame was reinserted in the frame list passed as parameter to qc_do_build_pkt(). Then <frm_list> was spliced to the packet frame list even its frames were not built, nor attached to any packet. Such frames had their ->pkt member set to NULL, but considered as built, then sent leading to a crash in qc_release_frm() where ->pkt is dereferenced. This issue was again reported by useful traces provided by Tristan in GH #1808. Must be backported to 2.6.	2022-08-27 18:33:19 +02:00
Frédéric Lécaille	e35463c767	BUG/MINOR: quic: Null packet dereferencing from qc_dup_pkt_frms() trace This function must duplicate frames be resent from packets. Some of them are still in flight, others have already been detected as lost. In this case the original frame ->pkt member is NULL. Add a trace to distinguish these cases. Thank you to Tristan for having reported this issue in GH #1808. Must be backported to 2.6.	2022-08-27 10:29:30 +02:00
William Lallemand	d78dfe7891	BUG/MINOR: httpclient: fix resolution with port Fix the resolution in the httpclient when a port is associated to a domain. The do-resolve action doesn't support a port in its input. Must be backported to 2.6. Require the "host_only" converter to be backported.	2022-08-26 17:00:22 +02:00
William Lallemand	dd754cba16	MINOR: sample: add the host_only and port_only converters Add 2 converters that can manipulate the value of an Host header. host_only will return the host without any port, and port_only will return the port.	2022-08-26 17:00:22 +02:00
Frédéric Lécaille	eba9088a7c	Revert "MINOR: quic: Remove useless traces about references to TX packets" This reverts commit `f61398a7ca`. After having checked a version with more traces and reproduced the issue as reported by Tristan in GH #1808, there are remaining cases where a duplicated but not already sent frame have to be marked as acked because the frame it was copied from was acknowledeged before its copied was sent. Must be backported to 2.6.	2022-08-25 16:06:48 +02:00
Frédéric Lécaille	f61398a7ca	MINOR: quic: Remove useless traces about references to TX packets Since this commit: "BUG/MINOR: quic: Wrong list_for_each_entry() use when building packets from qc_do_build_pkt()" there is no more reason that frames can be released without having been sent, i.e. frames with non null ->pkt member. This ->pkt is the packet the frame is attached to. Must be backported to 2.6.	2022-08-25 07:35:47 +02:00
Frédéric Lécaille	560ddfa003	CLEANUP: quic: Remove a useless check in qc_lstnr_pkt_rcv() This function parses the QUIC packet from a UDP datagram. It was originally supposed to be run by several thread. Here we remove a section of code where the current thread checks there is not another thread which has already inserted the new quic_conn it is trying to insert in the connections tree. Must be backported to 2.6 to ease the future backports to come.	2022-08-24 18:59:23 +02:00
Frédéric Lécaille	15773f2101	BUG/MINOR: quic: Stalled connections (missing I/O handler wakeup) This was due to a missing I/O handler tasklet wakeup in process_timer() when detecting packet loss. As, qc_release_lost_pkts() could remove the lost packets from the in flight packets count, qc_set_timer() could cancel the timer used to wakeup the connection I/O handler. Then the connection could remain idle until it ends. Must be backported to 2.6.	2022-08-24 18:13:30 +02:00
Frédéric Lécaille	277c4629e7	BUG/MINOR: quic: Leak in qc_release_lost_pkts() for non in flight TX packets Packets with null "in flight" lengths are kept as the others packets as sent but not already acknowledeged in the by packet number space trees. But qc_release_lost_pkts() relied on this in fligh length to release the memory allocated for this packets. We must release the memory allocated for all the lost packets regardless of their in fligh lengths. Modify this function to do nothing if the list of lost packets passed as argument is empty. Stop using <lost_bytes> variable to decide if some packets memory must be released or not. Modify the callers to stop checking if this list is empty. Should helping in fixing memory leak as reported by Tristan in GH #1801. Must be backported to 2.6.	2022-08-24 18:13:30 +02:00
Frédéric Lécaille	5f6c25e447	Revert "BUG/MINOR: quix: Memleak for non in flight TX packets" This reverts commit `da9c441886`. Indeed this commit prevented the ACK only packets to be used as other packets when they are acknowledged. Even if not ack-eliciting packets they are acknowledged alongside others packets. Such acknowledged ACK only packets must be used for instance to compute the RTT. Must be backported to 2.6 if `da9c441` was backported to 2.6.	2022-08-24 18:12:59 +02:00
William Lallemand	b10b1196b8	MINOR: resolvers: shut the warning when "default" resolvers is implicit Shut the connect() warning of resolvers_finalize_config() when the configuration was not emitted manually. This shuts the warning for the "default" resolvers which is created automatically for the httpclient. Must be backported in 2.6.	2022-08-24 14:56:42 +02:00
Christopher Faulet	871dd82117	BUG/MINOR: tcpcheck: Disable QUICKACK only if data should be sent after connect It is only a real problem for agent-checks when there is no agent string to send. The condition to disable TCP_QUICKACK was only based on the action type following the connect one. But it is not always accurate. indeed, for agent-checks, there is always a SEND action. But if there is no "agent-send" string defined, nothing is sent. In this case, this adds 200ms of latency with no reason. To fix the bug, a flag is now used on the CONNECT action to instruct there are data that should be sent after the connect. For health-checks, this flag is set if the action following the connect is a SEND action. For agent-checks, it is set if an "agent-send" string is defined. This patch should fix the issue #1836. It must be backported as far as 2.2.	2022-08-24 11:59:04 +02:00
William Lallemand	6020c4e44e	BUG/MINOR: mworker: does not create the "default" resolvers in wait mode When doing a re-exec, the master was creating a "default" resolvers, which could result in a warning emitted because the "default" resolvers of the configuration file is not available anymore. Skip the creating of the "default" resolvers in wait mode, this is not useful in the master. Must be backported as far as 2.6.	2022-08-24 11:28:29 +02:00
William Lallemand	866b88bc95	BUG/MINOR: resolvers: return the correct value in resolvers_finalize_config() Patch `c31577f` ("MEDIUM: resolvers: continue startup if network is unavailable") was not working correctly. Indeed resolvers_finalize_config() was returning a ERR type, but a postparser is supposed to return 0 or 1. The return value was never right, however it was only a problem since `c31577f`. Must be backported in every stable branch.	2022-08-24 10:11:17 +02:00
Brad Smith	02fd3caa8f	BUILD: tcp_sample: fix build of get_tcp_info() on OpenBSD The build on OpenBSD is broken since commit `5c83e3a15` ("MINOR: tcp_sample: clarifying samples support per os, for further expansion."), hence it only affects 2.7 and 2.6. It looks like this changed things in such a way that if TCP_INFO is added but the OS is not added to the list of OS's it will not build. Extend support for get_tcp_info to OpenBSD. This must be backported to 2.6.	2022-08-24 05:23:13 +02:00
Willy Tarreau	8bd146d8af	MEDIUM: peers: limit the number of updates sent at once As seen in GH issue #1770, peers synchronization do not cope well with very large buffers because by default the only two reasons for stopping the processing of updates is either that the end was reached or that the buffer is full. This can cause high latencies, and even rightfully trigger the watchdog when the operations are numerous and slowed down by competition on the stick-table lock. This patch introduces a limit to the number of messages one may send at once, which now defaults to 200, regardless of the buffer size. This means taking and releasing the lock up to 400 times in a row, which is costly enough to let some other parts work. After some observation this could be backported to 2.6. If so, however, previous commits "BUG/MEDIUM: applet: fix incorrect check for abnormal return condition from handler" and "BUG/MINOR: applet: make the call_rate only count the no-progress calls" must be backported otherwise the call rate might trigger the looping protection.	2022-08-23 20:19:11 +02:00
Willy Tarreau	df3cab1ca1	BUG/MINOR: applet: make the call_rate only count the no-progress calls This is very similar to what we did in commit `6c539c4b8` ("BUG/MINOR: stream: make the call_rate only count the no-progress calls"), it's better to only count the call rate with no progress than to count all calls and try to figure if there's no progress, because a fast running applet might once satisfy the whole condition and trigger the bug. This typically happens when artificially limiting the number of messages sent at once by an applet, but could happen with plenty of highly interactive applets. This patch could be backported to stable versions if there are any indications that it might be useful there.	2022-08-23 20:19:11 +02:00
Willy Tarreau	8a3f58280f	BUG/MEDIUM: applet: fix incorrect check for abnormal return condition from handler We have quite numerous checks for abnormal applet handler behavior which are supposed to trigger the loop protection. However, consecutive to commit `15252cd9c` ("MEDIUM: stconn: move the RXBLK flags to the stream connector") that was merged into 2.6-dev12, one flag was incorrectly renamed, and the check for an applet waiting for a buffer that is present mistakenly turned to a check for missing room in the buffer. This erroneous test could mistakenly trigger on applets that perform intensive I/Os doing small exchanges each (e.g. cache, peers or HTTP client) if the load would be sustained (>100k iops). For the cache this could represent higher than 13 Gbps on an object at least 1.6 GB large for example, which is quite unlikely but theoretically possible. This fix needs to be backported to 2.6.	2022-08-23 20:19:11 +02:00
Frédéric Lécaille	a2d8ad20a3	MINOR: quic: Replace MT_LISTs by LISTs for RX packets. Replace ->rx.pqpkts quic_enc_level struct member MT_LIST by an LIST. Same thing for ->list quic_rx_packet struct member MT_LIST. Update the code consequently. This was a reminisence of the multithreading support (several threads by connection). Must be backported to 2.6	2022-08-23 17:55:02 +02:00
Frédéric Lécaille	b8047de11a	BUG/MINOR: quic: Safer QUIC frame builders Do not rely on the fact the callers of qc_build_frm() handle their buffer passed to function the correct way (without leaving garbage). Make qc_build_frm() update the buffer passed as argument only if the frame it builds is well formed. As far as I sse, there is no such callers which does not handle carefully such buffers. Must be backported to 2.6.	2022-08-23 17:40:09 +02:00
Frédéric Lécaille	a8a6043240	BUG/MINOR: quic: Wrong list_for_each_entry() use when building packets from qc_do_build_pkt() This is list_for_each_entry_safe() which must be used if we want to delete elements inside its code block. This could explain that some frames which were not built were added to packets with a NULL ->pkt member. Thank you to Tristan for having reported this issue through backtraces in GH #1808 Must be backported to 2.6.	2022-08-23 12:06:40 +02:00
Frédéric Lécaille	da9c441886	BUG/MINOR: quix: Memleak for non in flight TX packets First, these packets must not be inserted in the tree of TX packets. They are never explicitely acknowledged (for instance an ACK only packet will never be acknowledged). Furthermore, if taken into an account these packets may uselessly disturb the congestion control. We do not care if they are lost or not. Furthermore as the ->in_fligh_len member value is null they were not released by qc_release_lost_pkts() which rely on these values to decide to release the allocated memory for such packets. Must be backported to 2.6.	2022-08-22 19:06:08 +02:00
Emeric Brun	8032a276ce	BUG/MAJOR: mworker: fix infinite loop on master with no proxies. The master is re-exec with an empty proxies list if no master CLI is configured. This results in infinite loop since last patch: `3b68b602` ("BUG/MAJOR: log-forward: Fix log-forward proxies not fully initialized") This patch avoid to loop again on log-forward proxies list if empty. This patch should be backported until v2.3	2022-08-22 13:09:29 +02:00
Willy Tarreau	f1cfd9bc97	MINOR: cpu-map: remove obsolete diag warning about combined ranges We used to emit a diag warning in case ranges were used both with the process and thread part of a thread spec. Now with groups it's not longer a problem, so let's just kill this warning.	2022-08-22 10:46:13 +02:00
Willy Tarreau	3cd71acd06	BUG/MEDIUM: cpu-map: fix thread 1's affinity affecting all threads Since 2.7-dev2 with commit `5b09341c02` ("MEDIUM: cpu-map: replace the process number with the thread group number"), the thread group has replaced the process number in the "cpu-map" directive. In part due to a design limit in 2.4 and 2.5, a special case was made of thread 1 in commit `bda7c1decd` ("MEDIUM: config: simplify cpu-map handling"), because there was no other location to store a single-threaded setup's mask by then. The combination of the two resulted in a problem with thread groups, by which as soon as one line exhibiting thread number 1 alone was found in a config, the mask would be applied to all threads in the group. The loop was reworked to avoid this obsolete special case, and was factored for better legibility. One obsolete comment about nbproc was also removed. No backport is needed.	2022-08-22 10:38:00 +02:00
Frédéric Lécaille	ea4a5cbbdf	BUG/MINOR: mux-quic: Fix memleak on QUIC stream buffer for unacknowledged data Some clients send CONNECTION_CLOSE frame without acknowledging the STREAM data haproxy has sent. In this case, when closing the connection if there were remaining data in QUIC stream buffers, they were not released. Add a <closing> boolean option to qc_stream_desc_free() to force the stream buffer memory releasing upon closing connection. Thank you to Tristan for having reported such a memory leak issue in GH #1801. Must be backported to 2.6.	2022-08-20 19:08:31 +02:00
William Lallemand	62c0b99e3b	MINOR: ssl/cli: implement "add ssl ca-file" In ticket #1805 an user is impacted by the limitation of size of the CLI buffer when updating a ca-file. This patch allows a user to append new certificates to a ca-file instead of trying to put them all with "set ssl ca-file" The implementation use a new function ssl_store_dup_cafile_entry() which duplicates a cafile_entry and its X509_STORE. ssl_store_load_ca_from_buf() was modified to take an apped parameter so we could share the function for "set" and "add".	2022-08-19 19:58:53 +02:00
William Lallemand	d4774d3cfa	MINOR: ssl: handle ca-file appending in cafile_entry In order to be able to append new CA in a cafile_entry, ssl_store_load_ca_from_buf() was reworked and a "append" parameter was added. The function is able to keep the previous X509_STORE which was already present in the cafile_entry.	2022-08-19 19:58:53 +02:00
William Lallemand	ec7eb59d20	BUG/MINOR: ssl/cli: error when the ca-file is empty "set ssl ca-file" does not return any error when a ca-file is empty or only contains comments. This could be a problem is the file was malformated and did not contain any PEM header. It must be backported as far as 2.5.	2022-08-19 19:56:53 +02:00
Frédéric Lécaille	86a53c5669	MINOR: quic: Add reusable cipher contexts for header protection Implement quic_tls_rx_hp_ctx_init() and quic_tls_tx_hp_ctx_init() to initiliaze such header protection cipher contexts for each RX and TX parts and for each packet number spaces, only one time by connection. Make qc_new_isecs() call these two functions to initialize the cipher contexts of the Initial secrets. Same thing for ha_quic_set_encryption_secrets() to initialize the cipher contexts of the subsequent derived secrets (ORTT, 1RTT, Handshake). Modify qc_do_rm_hp() and quic_apply_header_protection() to reuse these cipher contexts. Note that there is no need to modify the key update for the header protection. The header protection secrets are never updated.	2022-08-19 18:31:59 +02:00
Emeric Brun	a8942cd9c4	BUG/MAJOR: log-forward: Fix ssl layer not initialized on bind even if configured Since commit `2071a99df` ("MINOR: listener/ssl: set the SSL xprt layer only once the whole config is known") the xprt is initialized for ssl directly from a generic funtion used to parse bind args. But the 'bind' lines from 'log-forward' sections were forgotten in commit `55f0f7bb5` ("MINOR: config: use the new bind_parse_args_list() to parse a "bind" line"). This patch re-works 'log-forward' section parsing to use the generic function to parse bind args and fix the issue. Since the generic way to parse was introduced in 2.6, this patch should be backported as far as this version.	2022-08-19 16:09:06 +02:00
Emeric Brun	3b68b60261	BUG/MAJOR: log-forward: Fix log-forward proxies not fully initialized Some initialisation for log forward proxies was missing such as ssl configuration on 'log-forward's 'bind' lines. After the loop on the proxy initialization code for proxies present in the main proxies list, this patch force to loop again on this code for proxies present in the log forward proxies list. Those two lists should be merged. This will be part of a global re-work of proxy initialization including peers proxies and resolver proxies. This patch was made in first attempt to fix the bug and to facilitate the backport on older branches waiting for a cleaner re-work on proxies initialization on the dev branch. This patch should be backported as far as 2.3.	2022-08-19 16:08:03 +02:00
Frédéric Lécaille	a846a17fde	MINOR: quic: Trace fix in qc_release_frm() This wrong trace came with this commit: "BUG/MINOR: quic: Possible crashes when dereferencing ->pkt quic_frame struct member" In qc_release_frm() we mark frames as acked. Nothing to see with references to frames. Thank you to Willy for having caught this one. Must be backported to 2.6 as these traces arrived with a bug fix to be backported to 2.6.	2022-08-19 12:15:05 +02:00
Frédéric Lécaille	e4c3074c00	MINOR: quic: Add the QUIC connection to mux traces This should help for debugging purpose. Should be backported to 2.6	2022-08-19 12:02:29 +02:00
Frédéric Lécaille	b827840b42	BUG/MINOR: quic: Wrong splitted duplicated frames handling When duplicated frames are splitted, we must propagate this information to the new allocated frame and add a reference to this new frame to the reference list of the original frame. Must be backported to 2.6	2022-08-19 10:10:43 +02:00
Frédéric Lécaille	2f16348d24	MINOR: quic: Add frame addresses to QUIC_EV_CONN_PRSAFRM event traces This should be useful to diagnose some issues. Should be backported to 2.6.	2022-08-19 09:59:07 +02:00
Frédéric Lécaille	1ba25c244e	BUG/MINOR: quic: Possible crashes when dereferencing ->pkt quic_frame struct member This was done at several places. First in qc_requeue_nacked_pkt_tx_frms. This aim of this function is, if needed, to requeue all the TX frames of a lost <pkt> packet passed as argument and detach them from this packet they have been sent from. They are possible cases where the frm->pkt quic_frame struct member could be NULL, as a result of a duplication of an original frame by qc_dup_pkt_frms(). This function adds the duplicated frame to the original frame reference list: LIST_APPEND(&origin->reflist, &dup_frm->ref); But, in this function, the packet which contains the frame is the one which is passed as argument (for debug purpose). So let us prefer using this variable. Also do not dereference this ->pkt quic_frame member in qc_release_frm() and qc_frm_unref() and add a trace to catch the frame with a null ->pkt member. They are logically frames which have not already been sent. Thank you to Tristan for having reported such crashes in GH #1808. Must be backported to 2.6	2022-08-19 09:58:28 +02:00
Willy Tarreau	473e0e54f5	BUG/MINOR: mux-h2: send a CANCEL instead of ES on truncated writes If a POST upload is cancelled after having advertised a content-length, or a response body is truncated after a content-length, we're not allowed to send ES because in this case the total body length must exactly match the advertised value. Till now that's what we were doing, and that was causing the other side (possibly haproxy) to respond with an RST_STREAM PROTOCOL_ERROR due to "ES on DATA frame before content-length". We can behave a bit cleaner here. Let's detect that we haven't sent everything, and send an RST_STREAM(CANCEL) instead, which is designed exactly for this purpose. This patch could be backported to older versions but only a little bit of exposure to make sure it doesn't wake up a bad behavior somewhere. It relies on the following previous commit: "MINOR: mux-h2: make streams know if they need to send more data"	2022-08-19 08:03:53 +02:00
Willy Tarreau	4877045f1d	MINOR: mux-h2: make streams know if they need to send more data H2 streams do not even know if they are expected to send more data or not, which is problematic when closing because we don't know if we're closing too early or not. Let's start by adding a new stream flag "H2_SF_MORE_HTX_DATA" to indicate this on the tx path.	2022-08-19 08:03:53 +02:00
Willy Tarreau	ed2b9d9f27	MINOR: mux-h2/traces: report transition to SETTINGS1 before not after Traces indicating "switching to XXX" generally apply before the transition so that the current connection state is visible in the trace. SETTINGS1 was incorrect in this regard, with the trace being emitted after. Let's fix this. No need to backport this, as this is purely cosmetic.	2022-08-19 08:03:53 +02:00
Willy Tarreau	0f45871344	BUG/MEDIUM: mux-h2: do not fiddle with ->dsi to indicate demux is idle When switching to H2_CS_FRAME_H, we do not want to present the previous frame's state, flags, length etc in traces, or we risk to confuse the analysis, making the reader think that the header information presented is related to the new frame header being analysed. A naive approach could have consisted in simply relying on the current parser state (FRAME_H being that state), but traces are emitted before switching the state, so traces cannot rely on this. This was initially addressed by commit `73db434f7` ("MINOR: h2/trace: report the frame type when known") which used to set dsi to -1 when the connection becomes idle again, but was accidentally broken by commit `5112a603d` ("BUG/MAJOR: mux_h2: Don't consume more payload than received for skipped frames") which moved dsi after calling the trace function. But in both cases there's problem with this approach. If an RST or WU frame cannot be uploaded due to a busy mux, and at the same time we complete processing on a perfect end of frame with no single new frame header, we can leave the demux loop with dsi=-1 and with RST or WU to be sent, and these ones will be sent for stream ID -1. This is what was reported in github issue #1830. This can be reproduced with a config chaining an h1->h2 proxy to an empty h2 frontend, and uploading a large body such as below: $ (printf "POST / HTTP/1.1\r\nContent-length: 1000000000\r\n\r\n"; cat /dev/zero) \| nc 0 4445 > /dev/null This shows that we must never affect ->dsi which must always remain valid, and instead we should set "something else". That something else could be served by the demux frame type, but that one also needs to be preserved for the RST_STREAM case. Instead, let's just add a connection flag to say that the demuxing is in progress. This will be set once a new demux header is set and reset after the end of a frame. This way the trace subsystem can know that dft/dfl must not be displayed, without affecting the logic relying on such values. Given that the commits above are old and were backported to 1.8, this new one also needs to be backported as far as 1.8. Many thanks to David le Blanc (@systemmonkey42) for spotting, reporting, capturing and analyzing this bug; his work permitted to quickly spot the problem.	2022-08-19 08:03:53 +02:00
Willy Tarreau	1addf8b777	BUG/MEDIUM: cli: always reset the service context between commands Erwan Le Goas reported that chaining certain commands on the CLI would systematically crash the process; for example, "show version; show sess". This happened since the conversion of cli context to appctx->svcctx, because if applet_reserve_svcctx() is called a first time for a tiny context, it's allocated in-situ, and later a keyword that wants a larger one will see that it's not null and will reuse it and will overwrite the end of the first one's context. What is missing is a reset of the svcctx when looping back to CLI_ST_GETREQ. This needs to be backported to 2.6, and relies on previous commit "MINOR: applet: add a function to reset the svcctx of an applet".	2022-08-18 18:16:36 +02:00
Willy Tarreau	1cc08a33e1	MINOR: applet: add a function to reset the svcctx of an applet The CLI needs to reset the svcctx between commands, and there was nothing done to handle this. Let's add appctx_reset_svcctx() to do that, it's the closing equivalent of appctx_reserve_svcctx(). This will have to be backported to 2.6 as it will be used by a subsequent patch to fix a bug.	2022-08-18 18:16:36 +02:00
Amaury Denoyelle	115ccce867	MEDIUM: h3: concatenate multiple cookie headers As specified by RFC 9114, multiple cookie headers must be concatenated into a single entry before passing it to a HTTP/1.1 connection. To implement this, reuse the same function as already used for HTTP/2 module. This should answer to feature requested in github issue #1818.	2022-08-18 16:13:33 +02:00
Amaury Denoyelle	2c5a7ee333	REORG: h2: extract cookies concat function in http_htx As specified by RFC 7540, multiple cookie headers are merged in a single entry before passing it to a HTTP/1.1 connection. This step is implemented during headers parsing in h2 module. Extract this code in the generic http_htx module. This will allow to reuse it quickly for HTTP/3 implementation which has the same requirement for cookie headers.	2022-08-18 16:13:33 +02:00
Amaury Denoyelle	704675656b	BUG/MEDIUM: quic: fix crash on MUX send notification MUX notification on TX has been edited recently : it will be notified only when sending its own data, and not for example on retransmission by the quic-conn layer. This is subject of the patch : `b29a1dc2f4` BUG/MINOR: quic: do not notify MUX on frame retransmit A new flag QUIC_FL_CONN_RETRANS_LOST_DATA has been introduced to differentiate qc_send_app_pkts invocation by MUX and directly by the quic-conn layer in quic_conn_app_io_cb(). However, this is a first problem as internal quic-conn layer usage is not limited to retransmission. For example for NEW_CONNECTION_ID emission. Another problem much important is that send functions are also called through quic_conn_io_cb() which has not been protected from MUX notification. This could probably result in crash when trying to notify the MUX. To fix both problems, quic-conn flagging has been inverted : when used by the MUX, quic-conn is flagged with QUIC_FL_CONN_TX_MUX_CONTEXT. To improve the API, MUX must now used qc_send_mux which ensure the flag is set. qc_send_app_pkts is now static and can only be used by the quic-conn layer. This must be backported wherever the previously mentionned patch is.	2022-08-18 11:33:22 +02:00
Frédéric Lécaille	4173a39c1f	BUG/MINOR: quic: Missing initializations for ducplicated frames. When duplication frames in qc_dup_pkt_frms(), ->pkt member was not correctly initialized (copied from the original frame). This could not have any impact because this member is initialized whe the frame is added to a packet. This was also the case for ->flags. Also replace the pool_zalloc() call by a call to pool_alloc(). Must be backported to 2.6.	2022-08-18 10:28:31 +02:00
Mateusz Malek	4b85a963be	BUG/MEDIUM: http-ana: fix crash or wrong header deletion by http-restrict-req-hdr-names When using `option http-restrict-req-hdr-names delete`, HAproxy may crash or delete wrong header after receiving request containing multiple forbidden characters in single header name; exact behavior depends on number of request headers, number of forbidden characters and position of header containing them. This patch fixes GitHub issue #1822. Must be backported as far as 2.2 (buggy feature got included in 2.2.25, 2.4.18 and 2.5.8).	2022-08-17 15:52:17 +02:00
Amaury Denoyelle	b29a1dc2f4	BUG/MINOR: quic: do not notify MUX on frame retransmit On STREAM emission, quic-conn notifies MUX through a callback named qcc_streams_sent_done(). This also happens on retransmission : in this case offset are examined and notification is ignored if already seen. However, this behavior has slightly changed since `e53b489826` BUG/MEDIUM: mux-quic: fix server chunked encoding response Indeed, if offset diff is NULL, frame is now not ignored. This is to support FIN notification with a final empty STREAM frame. A side-effect of this is that if the last stream frame is retransmitted, it won't be ignored in qcc_streams_sent_done(). In most cases, this side-effect is harmless as qcs instance will soon be freed after being closed. But if qcs is still alive, this will cause a BUG_ON crash as it is considered as locally closed. This bug depends on delay condition and seems to be extremely rare. But it might be the reason for a crash seen on interop with s2n client on http3 testcase : FATAL: bug condition "qcs->st == QC_SS_CLO" matched at src/mux_quic.c:372 call trace(16): \| 0x558228912b0d [b8 01 00 00 00 c6 00 00]: main-0x1c7878 \| 0x558228917a70 [48 8b 55 d8 48 8b 45 e0]: qcc_streams_sent_done+0xcf/0x355 \| 0x558228906ff1 [e9 29 05 00 00 48 8b 05]: main-0x1d3394 \| 0x558228907cd9 [48 83 c4 10 85 c0 0f 85]: main-0x1d26ac \| 0x5582289089c1 [48 83 c4 50 85 c0 75 12]: main-0x1d19c4 \| 0x5582288f8d2a [48 83 c4 40 48 89 45 a0]: main-0x1e165b \| 0x5582288fc4cc [89 45 b4 83 7d b4 ff 74]: qc_send_app_pkts+0xc6/0x1f0 \| 0x5582288fd311 [85 c0 74 12 eb 01 90 48]: main-0x1dd074 \| 0x558228b2e4c1 [48 c7 c0 d0 60 ff ff 64]: run_tasks_from_lists+0x4e6/0x98e \| 0x558228b2f13f [8b 55 80 29 c2 89 d0 89]: process_runnable_tasks+0x7d6/0x84c \| 0x558228ad9aa9 [8b 05 75 16 4b 00 83 f8]: run_poll_loop+0x80/0x48c \| 0x558228ada12f [48 8b 05 aa c5 20 00 48]: main-0x256 \| 0x7ff01ed2e609 [64 48 89 04 25 30 06 00]: libpthread:+0x8609 \| 0x7ff01e8ca163 [48 89 c7 b8 3c 00 00 00]: libc:clone+0x43/0x5e To reproduce it locally, code was artificially patched to produce retransmission and avoid qcs liberation. In order to fix this and avoid future class of similar problem, the best way is to not call qcc_streams_sent_done() to notify MUX for retranmission. To implement this, we test if any of QUIC_FL_CONN_RETRANS_OLD_DATA or the new flag QUIC_FL_CONN_RETRANS_LOST_DATA is set. A new wrapper qc_send_app_retransmit() has been added to set the new flag as a complement to already existing qc_send_app_probing(). This must be backported up to 2.6.	2022-08-17 11:06:24 +02:00
Amaury Denoyelle	cc13047364	MINOR: quic: refactor application send Adjust qc_send_app_pkts function : remove <old_data> arg and provide a new wrapper function qc_send_app_probing() which should be used instead when probing with old data. This simplifies the interface of the default function, most notably for the MUX which does not interfer with retransmission. QUIC_FL_CONN_RETRANS_OLD_DATA flag is set/unset directly in the wrapper qc_send_app_probing(). At the same time, function documentation has been updated to clarified arguments and return values. This commit will be useful for the next patch to differentiate MUX and retransmission send context. As a consequence, the current patch should be backported wherever the next one will be.	2022-08-17 11:05:49 +02:00
Amaury Denoyelle	3baab744e7	MINOR: mux-quic: add missing args on some traces Complete some MUX traces by adding qcc or qcs instance as arguments when this is possible. This will be useful when several connections are interleaved.	2022-08-17 11:05:47 +02:00
Amaury Denoyelle	fd79ddb2d6	MINOR: mux-quic: adjust traces on stream init Adjust traces on qcc_init_stream_remote() : replace "opening" by "initializing" to avoid confusion with traces dealing with OPEN stream state.	2022-08-17 11:05:46 +02:00
Amaury Denoyelle	bf3c208760	BUG/MEDIUM: mux-quic: reject uni stream ID exceeding flow control Emit STREAM_LIMIT_ERROR if a client tries to open an unidirectional stream with an ID greater than the value specified by our flow-control limit. The code is similar to the bidirectional stream opening. MAX_STREAMS_UNI emission is not implement for the moment and is left as a TODO. This should not be too urgent for the moment : in HTTP/3, a client has only a limited use for unidirectional streams (H3 control stream + 2 QPACK streams). This is covered by the value provided by haproxy in transport parameters. This patch has been tagged with BUG as it should have prevented last crash reported on github issue #1808 when opening a new unidirectional streams with an invalid ID. However, it is probably not the main cause of the bug contrary to the patch commit `11a6f4007b` BUG/MINOR: quic: Wrong status returned by qc_pkt_decrypt() This must be backported up to 2.6.	2022-08-17 11:05:19 +02:00
Amaury Denoyelle	26aa399d6b	MINOR: qpack: report error on enc/dec stream close As specified by RFC 9204, encoder and decoder streams must not be closed. If the peer behaves incorrectly and closes one of them, emit a H3_CLOSED_CRITICAL_STREAM connection error. To implement this, QPACK stream decoding API has been slightly adjusted. Firstly, fin parameter is passed to notify about FIN STREAM bit. Secondly, qcs instance is passed via unused void* context. This allows to use qcc_emit_cc_app() function to report a CONNECTION_CLOSE error.	2022-08-17 11:04:53 +02:00
Amaury Denoyelle	6b02c6bb47	MINOR: h3: report error on control stream close As specified by RFC 9114 the control stream must not be closed. If the peer behaves incorrectly and closes it, emit a H3_CLOSED_CRITICAL_STREAM connection error.	2022-08-17 11:04:52 +02:00
Amaury Denoyelle	f372e744de	MINOR: quic: adjust quic_frame flag manipulation Replace a plain '=' operator by '\|=' when setting quic_frame QUIC_FL_TX_FRAME_LOST flag. For the moment, this change has no impact as only two exclusive flags are defined for quic_frame. On the edited code path we are certain that QUIC_FL_TX_FRAME_ACKED is not set due to a previous if statement, so a plain equal or a binary OR is strictly identical. This change will be useful if new flags are defined for quic_frame in the future. These new flags won't be resetted automatically thanks to binary OR without explictly intended, which otherwise could easily lead to new bugs.	2022-08-17 11:04:47 +02:00
Fr�d�ric L�caille	bbeec37b31	MINOR: stick-table: Add table_expire() and table_idle() new converters table_expire() returns the expiration delay for a stick-table entry associated to an input sample. Its counterpart table_idle() returns the time the entry remained idle since the last time it was updated. Both converters may take a default value as second argument which is returned when the entry is not present.	2022-08-17 10:52:15 +02:00
Willy Tarreau	cc1a2a1867	MINOR: chunk: inline alloc_trash_chunk() This function is responsible for all calls to pool_alloc(trash), whose total size can be huge. As such it's quite a pain that it doesn't provide more hints about its users. However, since the function is tiny, it fully makes sense to inline it, the code is less than 0.1% larger with this. This way we can now detect where the callers are via "show profiling", e.g.: 0 1953671 0 32071463136\| 0x59960f main+0x10676f p_free(-16416) [pool=trash] 0 1 0 16416\| 0x59960f main+0x10676f p_free(-16416) [pool=trash] 1953672 0 32071479552 0\| 0x599561 main+0x1066c1 p_alloc(16416) [pool=trash] 0 976835 0 16035723360\| 0x576ca7 http_reply_to_htx+0x447/0x920 p_free(-16416) [pool=trash] 0 1 0 16416\| 0x576ca7 http_reply_to_htx+0x447/0x920 p_free(-16416) [pool=trash] 976835 0 16035723360 0\| 0x576a5d http_reply_to_htx+0x1fd/0x920 p_alloc(16416) [pool=trash] 1 0 16416 0\| 0x576a5d http_reply_to_htx+0x1fd/0x920 p_alloc(16416) [pool=trash]	2022-08-17 10:45:22 +02:00
Willy Tarreau	42b180dcdb	MINOR: pools/memprof: store and report the pool's name in each bin Storing the pointer to the pool along with the stats is quite useful as it allows to report the name. That's what we're doing here. We could store it in place of another field but that's not convenient as it would require to change all functions that manipulate counters. Thus here we store one extra field, as well as some padding because the struct turns 56 bytes long, thus better go to 64 directly. Example of output from "show profiling memory": 2 0 48 0\| 0x4bfb2c ha_quic_set_encryption_secrets+0xcc/0xb5e p_alloc(24) [pool=quic_tls_iv] 0 55252 0 10608384\| 0x4bed32 main+0x2beb2 free(-192) 15 0 2760 0\| 0x4be855 main+0x2b9d5 p_alloc(184) [pool=quic_frame] 1 0 1048 0\| 0x4be266 ha_quic_add_handshake_data+0x2b6/0x66d p_alloc(1048) [pool=quic_crypto] 3 0 552 0\| 0x4be142 ha_quic_add_handshake_data+0x192/0x66d p_alloc(184) [pool=quic_frame] 31276 0 6755616 0\| 0x4bb8f9 quic_sock_fd_iocb+0x689/0x69b p_alloc(216) [pool=quic_dgram] 0 31424 0 6787584\| 0x4bb7f3 quic_sock_fd_iocb+0x583/0x69b p_free(-216) [pool=quic_dgram] 152 0 32832 0\| 0x4bb4d9 quic_sock_fd_iocb+0x269/0x69b p_alloc(216) [pool=quic_dgram]	2022-08-17 10:34:00 +02:00
Willy Tarreau	facfad2b64	MINOR: pool/memprof: report pool alloc/free in memory profiling Pools are being used so well that it becomes difficult to profile their usage via the regular memory profiling. Let's add new entries for pools there, named "p_alloc" and "p_free" that correspond to pool_alloc() and pool_free(). Ideally it would be nice to only report those that fail cache lookups but that's complicated, particularly on the free() path since free lists are released in clusters to the shared pools. It's worth noting that the alloc_tot/free_tot fields can easily be determined by multiplying alloc_calls/free_calls by the pool's size, and could be better used to store a pointer to the pool itself. However it would require significant changes down the code that sorts output. If this were to cause a measurable slowdown, an alternate approach could consist in using a different value of USE_MEMORY_PROFILING to enable pools profiling. Also, this profiler doesn't depend on intercepting regular malloc functions, so we could also imagine enabling it alone or the other one alone or both. Tests show that the CPU overhead on QUIC (which is already an extremely intensive user of pools) jumps from ~7% to ~10%. This is quite acceptable in most deployments.	2022-08-17 09:38:05 +02:00
Willy Tarreau	219afa2ca8	MINOR: memprof: export the minimum definitions for memory profiling Right now it's not possible to feed memory profiling info from outside activity.c, so let's export the function and move the enum and struct to the include file.	2022-08-17 09:03:57 +02:00
Frédéric Lécaille	11a6f4007b	BUG/MINOR: quic: Wrong status returned by qc_pkt_decrypt() This bug came with this big commit: "MEDIUM: quic: xprt traces rework" This is the <ret> variable value which must be returned by most of the xprt functions. This leaded packets which could not be decrypted to be parsed, with weird frames to be parsed as found by Tristan in GH #1808. To be backported where the commit above was backported.	2022-08-16 14:54:32 +02:00
Frédéric Lécaille	ebb1070721	BUG/MINOR: quic: MIssing check when building TX packets When building an ack-eliciting frame only packet, if we did not manage to add at least one such a frame to the packet, we did not notify the caller about the fact the packet is empty. This could lead the caller to believe everything was ok and make it endlessly try to build packet again and again. This issue was amplified by the recent changes where a while(1) loop has been added to qc_send_app_pkt() which calls qc_do_build_pkt() through qc_prep_app_pkts() until we could not prepare packets. Before this recent change, I guess only one empty packet was sent. This patch checks that non empty packets could be built by qc_do_build_pkt() and makes this function return an error if this was the case. Also note that such an issue could happened only when the packet building was limited by the congestion control. Thank you to Tristan for having reported this issue in GH #1808. Must be backported to 2.6.	2022-08-16 12:23:38 +02:00
Amaury Denoyelle	35a66c0a36	BUG/MINOR: mux-quic: fix crash with traces in qc_detach() qc_detach() is used to free a qcs as notified by sedesc. If there is no more stream active and the connection is considered as dead, it will then be freed. This prevent to dereference qcc in TRACE macro. Else this will cause a crash. Use a different code-path on release for qc_detach() to fix this bug. This will fix the last occurence of crash on github issue #1808. This has been introduced by recent QUIC MUX traces rework. Thus, it does not need to be backport.	2022-08-12 16:02:00 +02:00
Willy Tarreau	ded77cc71f	MINOR: ring: archive a previous file-backed ring on startup In order to ensure that an instant restart of the process will not wipe precious debugging information, and to leave time for an admin to archive a copy of a ring, now upon startup, any previously existing file will be renamed with the extra suffix ".bak", and any previously existing file with suffix ".bak" will be removed.	2022-08-12 15:40:19 +02:00
Willy Tarreau	8e87705c21	BUILD: sink: replace S_IRUSR, S_IWUSR with their octal value The build broke on freebsd with S_IRUSR undefined after commit `0b8e9ceb1` ("MINOR: ring: add support for a backing-file"). Maybe another include is needed there, but the point is that we really don't care about these symbolic names, file modes are more readable as 0600 than via these cryptic names anyway, so let's go back to 0600. This will teach me not to try to make things too clean. No backport is needed.	2022-08-12 15:03:12 +02:00
Fr�d�ric L�caille	bfb077acff	BUG/MINOR: quic: memleak on wrong datagram receipt There was a missing pool_free() call for such datagrams. As far as I see there is no leak on valid datagram receipt. Must be backported to 2.6.	2022-08-12 12:19:26 +02:00
Willy Tarreau	0b8e9ceb12	MINOR: ring: add support for a backing-file This mmaps a file which will serve as the backing-store for the ring's contents. The idea is to provide a way to retrieve sensitive information (last logs, debugging traces) even after the process stops and even after a possible crash. Right now this was possible by connecting to the CLI and dumping the contents of the ring live, but this is not handy and consumes quite a bit of resources before it is needed. With a backing file, the ring is effectively RAM-mapped file, so that contents stored there are the same as those found in the file (the OS doesn't guarantee immediate sync but if the process dies it will be OK). Note that doing that on a filesystem backed by a physical device is a bad idea, as it will induce slowdowns at high loads. It's really important that the device is RAM-based. Also, this may have security implications: if the file is corrupted by another process, the storage area could be corrupted, causing haproxy to crash or to overwrite its own memory. As such this should only be used for debugging.	2022-08-12 11:18:46 +02:00
Willy Tarreau	6df10d872b	MINOR: ring: support creating a ring from a linear area Instead of allocating two parts, one for the ring struct itself and one for the storage area, ring_make_from_area() will arrange the two inside the same memory area, with the storage starting immediately after the struct. This will allow to store a complete ring state in shared memory areas for example.	2022-08-12 11:18:46 +02:00
Fr�d�ric L�caille	7629f5d670	BUG/MEDIUM: quic: Wrong use of <token_odcid> in qc_lsntr_pkt_rcv() This commit was not complete: "BUG/MEDIUM: quic: Possible use of uninitialized <odcid> variable in qc_lstnr_params_init()" <token_odcid> should have been directly passed to qc_lstnr_params_init() without dereferencing it to prevent haproxy to have new chances to crash! Must be backported to 2.6.	2022-08-11 19:12:12 +02:00
Willy Tarreau	18d1306abd	BUG/MEDIUM: ring: fix too lax 'size' parser It took me a while to figure why a ring declared with "size 1M" was causing strange effects in a ring, it's just because it's parsed as "1", which is smaller than the default 16384 size and errors are silently ignored. This commit tries to address this the best possible way without breaking existing configs that would work by accident, by warning that the size is ignored if it's smaller than the current one, and by printing the parsed size instead of the input string in warnings and errors. This way if some users have "size 10000" or "size 100k" it will continue to work as 16kB like today but they will now be aware of it. In addition the error messages were a bit poor in context in that they only provided line numbers. The ring name was added to ease locating the problem. As the issue was present since day one and was introduced in 2.2 with commit `99c453df9d` ("MEDIUM: ring: new section ring to declare custom ring buffers."), it could make sense to backport this as far as 2.2, but with 2.2 being quite old now it doesn't seem very reasonable to start emitting new config warnings in config that apparently worked well. Thus it looks more reasonable to backport this as far as 2.4.	2022-08-11 19:05:19 +02:00
Fr�d�ric L�caille	e9325e97c2	BUG/MEDIUM: quic: Possible use of uninitialized <odcid> variable in qc_lstnr_params_init() When receiving a token into a client Initial packet without a cluster secret defined by configuration, the <odcid> variable used to parse the ODCID from the token could be used without having been initialized. Such a packet must be dropped. So the sufficient part of this patch is this check: + } + else if (!global.cluster_secret && token_len) { + /* Impossible case: a token was received without configured + * cluster secret. + */ + TRACE_PROTO("Packet dropped", QUIC_EV_CONN_LPKT, + NULL, NULL, NULL, qv); + goto drop; } Take the opportunity of this patch to rework and make it more readable this part of code where such a packet must be dropped removing the <check_token> variable. When an ODCID is parsed from a token, new <token_odcid> new pointer variable is set to the address of the parsed ODCID. This way, is not set but used it will make crash haproxy. This was not always the case with an uninitialized local variable. Adapt the API to used such a pointer variable: <token> boolean variable is removed from qc_lstnr_params_init() prototype. This must be backported to 2.6.	2022-08-11 18:33:36 +02:00
Amaury Denoyelle	6bdf9367fb	BUG/MEDIUM: mux-quic: fix crash due to invalid trace arg Traces argument were incorrectly used in qcs_free(). A qcs was specified as first arg instead of a connection. This will lead to a crash if developer qmux traces are activated. This is now fixed. This bug has been introduced with QUIC MUX traces rework. No need to backport.	2022-08-11 18:24:53 +02:00
Amaury Denoyelle	4c9a1642c1	MINOR: mux-quic: define new traces Add new traces to help debugging on QUIC MUX. Most notable, the following functions are now traced : * qcc_emit_cc * qcs_free * qcs_consume * qcc_decode_qcs * qcc_emit_cc_app * qcc_install_app_ops * qcc_release_remote_stream * qcc_streams_sent_done * qc_init	2022-08-11 15:20:44 +02:00
Amaury Denoyelle	047d86a34b	CLEANUP: mux-quic: adjust traces level Change default devel level for some traces in QUIC MUX: * proto : used to notify about reception/emission of frames * state : modification of internal state of connection or streams * data : detailled information about transfer and flow-control	2022-08-11 15:20:44 +02:00
Amaury Denoyelle	c7fb0d2b7a	MINOR: mux-quic: define protocol error traces Replace devel traces with error level on all errors situation. Also a new event QMUX_EV_PROTO_ERR is used. This should help to detect invalid situations quickly.	2022-08-11 15:20:44 +02:00
Amaury Denoyelle	f0b67f995c	MINOR: mux-quic: adjust enter/leave traces Improve MUX traces by adding some missing enter/leave trace points. In some places, early function returns have been replaced by a goto statement.	2022-08-11 15:20:43 +02:00
Fr�d�ric L�caille	59507de932	CLEANUP: quic: Remove trailing spaces This spaces have come with this commit: "MEDIUM: quic: xprt traces rework".	2022-08-11 14:33:43 +02:00
Fr�d�ric L�caille	96d08d37d9	BUG/MINOR: quic: Possible infinite loop in quic_build_post_handshake_frames() This loop is due to the fact that we do not select the next node before the conditional "continue" statement. Furthermore the condition and the "continue" statement may be removed after replacing eb64_first() call by eb64_lookup_ge(): we are sure this condition may not be satisfied. Add some comments: this function initializes connection IDs with sequence number 1 upto <max> non included. Take the opportunity of this patch to remove a "return" wich broke this traces rule: for any function, do not call TRACE_ENTER() without TRACE_LEAVE()! Add also TRACE_ERROR() for any encoutered errors. Must be backported to 2.6	2022-08-11 14:33:43 +02:00
Fr�d�ric L�caille	a6920a25d9	MINOR: quic: Remove useless lock for RX packets This lock was there be able to handle the RX packets for a connetion from several threads. This is no more needed since a QUIC connection is always handled by the same thread. May be backported to 2.6	2022-08-11 14:33:43 +02:00
Willy Tarreau	6a378d1677	BUILD: stconn: fix build warning at -O3 about possible null sc gcc-6.x and 7.x emit build warnings about sc possibly being null upon return from sc_detach_endp(). This actually is not the case and the compiler is a little bit overzealous there, but there exists code paths that can make this analysis non-trivial so let's at least add a similar BUG_ON() to let both the compiler and the deverloper know this doesn't happen. This should be backported to 2.6.	2022-08-11 13:59:13 +02:00
Fr�d�ric L�caille	a8b2f843d2	MEDIUM: quic: xprt traces rework Add a least as much as possible TRACE_ENTER() and TRACE_LEAVE() calls to any function. Note that some functions do not have any access to the a quic_conn argument when receiving or parsing datagram at very low level.	2022-08-11 11:11:20 +02:00
Willy Tarreau	e6ca435c04	BUG/MEDIUM: poller: use fd_delete() to release the poller pipes The poller pipes needed to communicate between multiple threads are allocated in init_pollers_per_thread() and released in deinit_pollers_per_thread(). The former adds them via fd_insert() so that they are known, but the former only closes them using a regular close(). This asymmetry represents a problem, because we have in the fdtab[] an entry for something that may disappear when one thread leaves, and since these FD numbers are very low, there is a very high likelihood that they are immediately reassigned to another thread trying to connect() to a server or just sending a health check. In this case, the other thread is going to fd_insert() the fd and the recently added consistency checks will notive that ->owner is not NULL and will crash. We just need to use fd_delete() here to match fd_insert(). Note that this test was added in 2.7-dev2 by commit `36d9097cf` ("MINOR: fd: Add BUG_ON checks on fd_insert()") which was backported to 2.4 as a safety measure (since it allowed to catch particularly serious issues). The patch in itself isn't wrong, it just revealed a long-dormant bug (been there since 1.9-dev1, 4 years ago). As such the current patch needs to be backported wherever the commit above is backported. Many thanks to Christian Ruppert for providing detailed traces in github issue #1807 and Cedric Paillet for bringing his complementary analysis that helped to understand the required conditions for this issue to happen (fast health checks @100ms + randomly long connections ~7s + fast reloads every second + hard-stop-after 5s were necessary on the dev's machine to trigger it from time to time).	2022-08-10 17:25:23 +02:00
Willy Tarreau	54bc78693d	BUG/MEDIUM: quic: always remove the connection from the accept list on close Fred managed to reproduce a crash showing a corrupted accept_list when firing thousands of concurrent picoquicdemo clients to a same instance. It may happen if the connection was placed into the accept_list and immediately closed before being processed (e.g. on error or t/o ?). In any case the quic_conn_release() function should always detach a connection to be deleted from any list, like it does for other lists, so let's add an MT_LIST_DELETE() here. This should be backported to 2.6.	2022-08-10 07:30:22 +02:00
Amaury Denoyelle	f0f92b2db8	BUG/MINOR: quic: fix crash on handshake io-cb for null next enc level When arriving at the handshake completion, next encryption level will be null on quic_conn_io_cb(). Thus this must be check this before dereferencing it via qc_need_sending() to prevent a crash. This was reproduced quickly when browsing over a local nextcloud instance through QUIC with firefox. This has been introduced in the current dev with quic-conn Tx refactoring. No need to backport it.	2022-08-09 18:01:10 +02:00
Amaury Denoyelle	96ca1b7c39	BUG/MINOR: mux-quic: open stream on STOP_SENDING Considered a stream as opened when receiving a STOP_SENDING frame as the first frame on the stream. This patch is tagged as BUG because a BUG_ON may occur if only a STOP_SENDING frame has been received for a frame. This will reset the stream in respect with RFC9000 but internally it is considered invalid transition to reset an idle stream. To fix this, simply use qcs_idle_open() on STOP_SENDING parsing function. This will mark the stream as OPEN before resetting it. This was detected on haproxy.org with the following backtrace : FATAL: bug condition "qcs->st == QC_SS_IDLE" matched at src/mux_quic.c:383 call trace(12): \| 0x490dd3 [b8 01 00 00 00 c6 00 00]: main-0x1d0633 \| 0x4975b8 [48 8b 85 58 ff ff ff 8b]: main-0x1c9e4e \| 0x497df4 [48 8b 45 c8 48 89 c7 e8]: main-0x1c9612 \| 0x49934c [48 8b 45 c8 48 89 c7 e8]: main-0x1c80ba \| 0x6b3475 [48 8b 05 54 1b 3a 00 64]: run_tasks_from_lists+0x45d/0x8b2 \| 0x6b4093 [29 c3 89 d8 89 45 d0 83]: process_runnable_tasks+0x7c9/0x824 \| 0x660bde [8b 05 fc b3 4f 00 83 f8]: run_poll_loop+0x74/0x430 \| 0x6611de [48 8b 05 7b a6 40 00 48]: main-0x228 \| 0x7f66e4fb2ea5 [64 48 89 04 25 30 06 00]: libpthread:+0x7ea5 \| 0x7f66e455ab0d [48 89 c7 e8 5b 72 fc ff]: libc:clone+0x6d/0x86 Stream states have been implemented in the current dev tree. Thus, this patch does not need to be backported.	2022-08-09 17:58:02 +02:00
Amaury Denoyelle	c09ef0c5fc	MINOR: quic: skip sending if no frame to send in io-cb Check on quic_conn_io_cb() if sending is required. This allows to skip over Tx buffer allocation if not needed. To implement this, we check if frame lists on current and next encryption level are empty. We also need to check if there is no need to send ACK, PROBE or CONNECTION_CLOSE. This has been isolated in a new function qc_need_sending() which may be reuse in some other functions in the future.	2022-08-09 16:03:49 +02:00
Amaury Denoyelle	654269c769	MINOR: quic: refactor datagram commit in Tx buffer This is the final patch on quic-conn Tx refactor. Extend the function which is used to write a datagram header to save at the same time written buffer data. This makes sense as the two operations are used at the same occasion when a pre-written datagram is comitted.	2022-08-09 16:00:30 +02:00
Amaury Denoyelle	5b68986d77	MINOR: quic: release Tx buffer on each send Complete refactor of quic-conn Tx buffer. The buffer is now released on every send operation completion. This should help to reduce memory footprint as now Tx buffers are allocated and released on demand. To simplify allocation/free of quic-conn Tx buffer, two static functions are created named qc_txb_alloc() and qc_txb_release().	2022-08-09 16:00:02 +02:00
Amaury Denoyelle	f2476053f9	MINOR: quic: replace custom buf on Tx by default struct buffer On first prototype version of QUIC, emission was multithreaded. To support this, a custom thread-safe ring-buffer has been implemented with qring/cbuf. Now the thread model has been adjusted : a quic-conn is always used on the same thread and emission is not multi-threaded. Thus, qring/cbuf usage can be replace by a standard struct buffer. The code has been simplified even more as for now buffer is always drained after a prepare/send invocation. This is the case since a datagram is always considered as sent even on sendto() error. BUG_ON statements guard are here to ensure that this model is always valid. Thus, code to handle data wrapping and consume too small contiguous space with a 0-length datagram is removed.	2022-08-09 15:45:47 +02:00
Amaury Denoyelle	56c6154dba	CLEANUP: mux-quic: remove loop on sending frames qc_send_app_pkts() has now a while loop implemented which allows to send all possible frames even if the send buffer is full between packet prepare and send. This is present since commit : `dc07751ed7` MINOR: quic: Send packets as much as possible from qc_send_app_pkts() This means we can remove code from the MUX which implement this at the upper layer. This is useful to simplify qc_send_frames() function. As mentionned commit is subject to backport, this commit should be backported as well to 2.6.	2022-08-09 15:41:07 +02:00
Willy Tarreau	4a426e2082	MINOR: debug/memstats: automatically determine first column size The first column's width may vary a lot depending on outputs, and it's annoying to have large empty columns on small names and mangled large columns that are not yet large enough. In order to overcome this, this patch adds a width field to the memstats applet's context, and this width is calculated the first time the function is entered, by estimating the width of all lines that will be dumped. This is simple enough and does the job well. If in the future some filtering criteria are added, it will still be possible to perform a single pass on everything depending on the desired output format.	2022-08-09 08:51:08 +02:00
Willy Tarreau	17200dd1f3	MINOR: debug: also store the function name in struct mem_stats The calling function name is now stored in the structure, and it's reported when the "all" argument is passed. The first column is significantly enlarged because some names are really wide :-(	2022-08-09 08:42:42 +02:00
Willy Tarreau	55c950baa9	MINOR: debug: store and report the pool's name in struct mem_stats Let's add a generic "extra" pointer to the struct mem_stats to store context-specific information. When tracing pool_alloc/pool_free, we can now store a pointer to the pool, which allows to report the pool name on an extra column. This significantly improves tracing capabilities. Example: proxy.c:1598 CALLOC size: 28832 calls: 4 size/call: 7208 dynbuf.c:55 P_FREE size: 32768 calls: 2 size/call: 16384 buffer quic_tls.h:385 P_FREE size: 34008 calls: 1417 size/call: 24 quic_tls_iv quic_tls.h:389 P_FREE size: 34008 calls: 1417 size/call: 24 quic_tls_iv quic_tls.h:554 P_FREE size: 34008 calls: 1417 size/call: 24 quic_tls_iv quic_tls.h:558 P_FREE size: 34008 calls: 1417 size/call: 24 quic_tls_iv quic_tls.h:562 P_FREE size: 34008 calls: 1417 size/call: 24 quic_tls_iv quic_tls.h:401 P_ALLOC size: 34080 calls: 1420 size/call: 24 quic_tls_iv quic_tls.h:403 P_ALLOC size: 34080 calls: 1420 size/call: 24 quic_tls_iv xprt_quic.c:4060 MALLOC size: 45376 calls: 5672 size/call: 8 quic_sock.c:328 P_ALLOC size: 46440 calls: 215 size/call: 216 quic_dgram	2022-08-09 08:26:59 +02:00
Frédéric Lécaille	ba19acd822	MINOR: quic: Replace pool_zalloc() by pool_malloc() for fake datagrams These fake datagrams are only used by the low level I/O handler. They are not provided to the "by connection" datagram handlers. This is why they are not MT_LIST_APPEND()ed to the listner RX buffer list (see &quic_dghdlrs[cid_tid].dgrams in quic_lstnr_dgram_dispatch(). Replace the call to pool_zalloc() to by the lighter call to pool_malloc() and initialize only the ->buf and ->length members. This is safe because only these fields are inspected by the low level I/O handler.	2022-08-08 21:10:58 +02:00
Frédéric Lécaille	ffde3168fc	BUG/MEDIUM: quic: Missing AEAD TAG check after removing header protection After removing the packet header protection, we can check the packet is long enough to contain a 16 bytes length AEAD TAG (at this end of the packet). This test was missing. Must be backported to 2.6.	2022-08-08 18:41:16 +02:00
Frédéric Lécaille	adc7641536	MINOR: quic: Too much useless traces in qc_build_frms() These traces about the available room into the packet currently built and its payload length could be displayed for each STREAM frame, even for those which have no chance to be embedded into a packet leading to very traces to be displayed from a connection with a lot of stream. This was revealed by traces provide by Tristan in GH #1808 May be backported to 2.6.	2022-08-08 16:18:55 +02:00
Frédéric Lécaille	99897d11d9	BUG/MEDIUM: quic: Wrong packet length check in qc_do_rm_hp() When entering this function, we first check the packet length is not too short. But this was done against the datagram lenght in place of the packet length. This could lead to the header protection to be removed using data past the end of the packet (without buffer overflow). Use the packet length in place of the datagram length which is at <end> address passed as parameter to this function. As the packet length is already stored in ->len packet struct member, this <end> parameter is no more useful. Must be backported to 2.6.	2022-08-08 11:02:04 +02:00
Willy Tarreau	2e64472d16	BUILD: cfgparse: always defined _GNU_SOURCE for sched.h and crypt.h _GNU_SOURCE used to be defined only when USE_LIBCRYPT was set. It's also needed for sched_setaffinity() to be exported. As a side effect, when USE_LIBCRYPT is not set, a warning is emitted, as Ilya found and reported in issue #1815. Let's just define _GNU_SOURCE regardless of USE_LIBCRYPT, and also explicitly add sched.h, as right now it appears to be inherited from one of the other includes. This should be backported to 2.4.	2022-08-07 16:55:07 +02:00
Ilya Shipitsin	52f2ff5b93	BUG/MEDIUM: fix DH length when EC key is used dh of length 1024 were chosen for EVP_PKEY_EC key type. let us pick "default_dh_param" instead. issue was found on Ubuntu 22.04 which is shipped with OpenSSL configured with SECLEVEL=2 by default. such SECLEVEL value prohibits DH shorter than 2048: OpenSSL error[0xa00018a] SSL_CTX_set0_tmp_dh_pkey: dh key too small better strategy for chosing DH still may be considered though.	2022-08-06 17:45:40 +02:00
Ilya Shipitsin	3b64a28e15	CLEANUP: assorted typo fixes in the code and comments This is 31st iteration of typo fixes	2022-08-06 17:12:51 +02:00
Willy Tarreau	c80bdb2da6	MINOR: threads: report the number of thread groups in build options haproxy -vv shows the number of threads but didn't report the number of groups, let's add it.	2022-08-06 16:45:26 +02:00
Willy Tarreau	f9d4a7dad3	BUG/MEDIUM: quic: break out of the loop in quic_lstnr_dghdlr The function processes packets sent by other threads in the current thread's queue. But if, for any reason, other threads write faster than the current one processes, this can lead to a situation where the function never returns. It seems that it might be what's happening in issue #1808, though unfortunately, this function is one of the rare without traces. But the amount of calls to functions like qc_lstnr_pkt_rcv() on a single thread seems to indicate this possibility. Thanks to Tristan for his efforts in collecting extremely precious traces! This likely needs to be backported to 2.6.	2022-08-05 16:12:00 +02:00
Amaury Denoyelle	6715cbf97f	BUG/MINOR: quic: adjust errno handling on sendto qc_snd_buf returned a size_t which means that it was never negative despite its documentation. Thus the caller who checked for this was never informed of a sendto error. Clean this by changing the return value of qc_snd_buf() to an integer. A 0 is returned on success. Every other values are considered as an error. This commit should be backported up to 2.6. Note that to not cause malfunctions, it must be backported after the previous patch : `906b058954` MINOR: quic: explicitely ignore sendto error This is to ensure that a sendto error does not cause send to be interrupted which may cause a stalled transfer without a proper retry mechanism. The impact of this bug seems null as caller explicitely ignores sendto error. However this part of code seems to be subject to strange issues and it may fix them in part. It may be of interest for github issue #1808.	2022-08-05 15:53:16 +02:00
Amaury Denoyelle	906b058954	MINOR: quic: explicitely ignore sendto error qc_snd_buf() returns an error if sendto has failed. On standard conditions, we should check for EAGAIN/EWOULDBLOCK errno and if so, register the file-descriptor in the poller to retry the operation later. However, quic_conn uses directly the listener fd which is shared for all QUIC connections of this listener on several threads. Thus, it's complicated to implement fd supversion via the poller : there is no mechanism to easily wakeup quic_conn or MUX after a sendto failure. A quick and simple solution for the moment is to considered a datagram as properly emitted even on sendto error. In the end, this will trigger the quic_conn retransmission timer as data will be considered lost on the network and the send operation will be retried. This solution will be replaced when fd management for quic_conn is reworked. In fact, this quick hack was already in use in the current code, albeit not voluntarily. This is due to a bug caused by an API mismatch on the return type of qc_snd_buf() which never emits a negative error code despite its documentation. Thus, all its invocation were considered as a success. If this bug was fixed, the sending would would have been interrupted by a break which could cause the transfer to freeze. qc_snd_buf() invocation is clean up : the break statement is removed. Send operation is now always explicitely conducted entirely even on error and buffer data is purged. A simple optimization has been added to skip over sendto when looping over several datagrams at the first sendto error. However, to properly function, it requires a fix on the return type of qc_snd_buf() which is provided in another patch. As the behavior before and after this patch seems identical, it is not labelled as a BUG. However, it should be backported for cleaning purpose. It may also have an impact on github issue #1808.	2022-08-05 15:45:25 +02:00
Frédéric Lécaille	e7df68a219	BUG/MINOR: quic: Missing Initial packet dropping case An Initial packet shorter than 1200 bytes must be dropped. The test was there without the "goto drop"! Must be backported to 2.6	2022-08-05 15:27:14 +02:00
Frédéric Lécaille	8ecb7363b5	MINOR: quic: Add two new stats counters for sendto() errors Add "quic_socket_full" new stats counter for sendto() errors with EAGAIN as errno. and "quic_sendto_err" counter for any other error.	2022-08-05 15:27:14 +02:00
Willy Tarreau	af5138fd07	BUG/MINOR: quic: do not reject datagrams matching minimum permitted size The dgram length check in quic_get_dgram_dcid() rejects datagrams matching exactly the minimum allowed length, which doesn't seem correct. I doubt any useful packet would be that small but better fix this to avoid confusing debugging sessions in the future. This might be backported to 2.6.	2022-08-05 10:31:29 +02:00
Willy Tarreau	53bfab080c	BUG/MINOR: sink: fix a race condition between the writer and the reader This is the same issue as just fixed in `b8e0fb97f` ("BUG/MINOR: ring/cli: fix a race condition between the writer and the reader") but this time for sinks. They're also sucking the ring and present the same race at high write loads. This must be backported to 2.2 as well. See comments in the aforementioned commit for backport hints if needed.	2022-08-04 17:21:16 +02:00
Christopher Faulet	96417f392d	BUG/MEDIUM: sink: Set the sink ref for forwarders created during ring parsing A reference to the sink was added in every forwarder by the commit `2ae25ea24` ("MINOR: sink: Add a ref to sink in the sink_forward_target structure"). But this commit is incomplete. It is not performed for the forwarders created during a ring parsing. This patch must be backported to 2.6.	2022-08-04 17:10:28 +02:00
Willy Tarreau	b8e0fb97f3	BUG/MINOR: ring/cli: fix a race condition between the writer and the reader The ring's CLI reader unlocks the read side of a ring and relocks it for writing only if it needs to re-subscribe. But during this time, the writer might have pushed data, see nobody subscribed hence woken nobody, while the reader would have left marking that the applet had no more data. This results in a dump that will not make any forward progress: the ring is clogged by this reader which believes there's no data and the writer will never wake it up. The right approach consists in verifying after re-attaching if the writer had made any progress in between, and to report that another call is needed. Note that a jump back to the beginning would also work but here we provide better fairness between readers this way. This needs to be backported to 2.2. The applet API needed to signal the availability of new data changed a few times since then.	2022-08-04 17:00:21 +02:00
Fr�d�ric L�caille	48bb875908	BUG/MINOR: quic: Avoid sending truncated datagrams There is a remaining loop in this ugly qc_snd_buf() function which could lead haproxy to send truncated UDP datagrams. For now on, we send a complete UDP datagram or nothing! Must be backported to 2.6.	2022-08-03 21:09:04 +02:00
Amaury Denoyelle	30e260e2e6	MEDIUM: mux-quic: implement http-request timeout Implement http-request timeout for QUIC MUX. It is used when the connection is opened and is triggered if no HTTP request is received in time. By HTTP request we mean at least a QUIC stream with a full header section. Then qcs instance is attached to a sedesc and upper layer is then responsible to wait for the rest of the request. This timeout is also used when new QUIC streams are opened during the connection lifetime to wait for full HTTP request on them. As it's possible to demux multiple streams in parallel with QUIC, each waiting stream is registered in a list <opening_list> stored in qcc with <start> as timestamp in qcs for the stream opening. Once a qcs is attached to a sedesc, it is removed from <opening_list>. When refreshing MUX timeout, if <opening_list> is not empty, the first waiting stream is used to set MUX timeout. This is efficient as streams are stored in the list in their creation order so CPU usage is minimal. Also, the size of the list is automatically restricted by flow control limitation so it should not grow too much. Streams are insert in <opening_list> by application protocol layer. This is because only application protocol can differentiate streams for HTTP messaging from internal usage. A function qcs_wait_http_req() has been added to register a request stream by app layer. QUIC MUX can then remove it from the list in qc_attach_sc(). As a side-note, it was necessary to implement attach qcc_app_ops callback on hq-interop module to be able to insert a stream in waiting list. Without this, a BUG_ON statement would be triggered when trying to remove the stream on sedesc attach. This is to ensure that every requests streams are registered for http-request timeout. MUX timeout is explicitely refreshed on MAX_STREAM_DATA and STOP_SENDING frame parsing to schedule http-request timeout if a new stream has been instantiated. It was already done on STREAM parsing due to a previous patch.	2022-08-03 15:04:18 +02:00
Amaury Denoyelle	6ec9837fca	MINOR: mux-quic: refactor refresh timeout function Try to reorganize qcc_refresh_timeout() to improve its readability. The main objective is to reduce the indentation level and if sequences by using goto statement to the end of the function. Also, backend and frontend code path should be more explicit with this new version.	2022-08-03 15:04:18 +02:00
Amaury Denoyelle	418ba21461	MINOR: mux-quic: refresh timeout on frame decoding Refresh the MUX connection timeout in frame parsing functions. This is necessary as these Rx operation are completed directly from the quic-conn layer outside of MUX I/O callback. Thus, the timeout should be refreshed on this occasion. Note that however on STREAM parsing refresh is only conducted when receiving the current consecutive data offset. Timeouts related function have been moved up in the source file to be able to use them in qcc_decode_qcs(). This commit will be useful for http-request timeout. Indeed, a new stream may be opened during qcc_decode_qcs() which should trigger this timeout until a full header section is received and qcs instance is attached to sedesc.	2022-08-03 15:04:18 +02:00
Amaury Denoyelle	8d818c6eab	MINOR: h3: support HTTP request framing state Store the current step of HTTP message in h3s stream. This reports if we are in the parsing of headers, content or trailers section. A new enum h3s_st_req is defined for this. This field is stored in h3s struct but only used for request stream. It is left undefined for other streams (control or QPACK streams). h3_is_frame_valid() has been extended to take into account this state information. A connection error H3_FRAME_UNEXPECTED is reported if an invalid frame according to the current state is received; for example a DATA frame at the beginning of a stream.	2022-08-03 15:04:18 +02:00
Frédéric Lécaille	2c77a5eb8e	BUG/MEDIUM: quic: Floating point exception in cubic_root() It is illegal to call my_flsl() with 0 as parameter value. It is a UB. This leaded cubic_root() to divide values by 0 at this line: x = 2 * x + (uint32_t)(val / ((uint64_t)x * (uint64_t)(x - 1))); Thank you to Tristan971 for having reported this issue in GH #1808 and Willy for having spotted the root cause of this bug. Must follow any cubic for QUIC backport (2.6).	2022-08-03 14:27:20 +02:00
Frédéric Lécaille	8ddde4f05e	BUG/MINOR: quic: Missing in flight ack eliciting packet counter decrement The decrement was missing in quic_pktns_tx_pkts_release() called each time a packet number space is discarded. This is not sure this bug could have an impact during handshakes. This counter is used to cancel the timer used both for packet detection and PTO, setting its value to null. So there could be retransmissions or probing which could be triggered for nothing. Must be backported to 2.6.	2022-08-03 12:59:59 +02:00
Christopher Faulet	6bb86539db	BUG/MEDIUM: proxy: Perform a custom copy for default server settings When a proxy is initialized with the settings of the default proxy, instead of doing a raw copy of the default server settings, a custom copy is now performed by calling srv_settings_copy(). This way, all settings will be really duplicated. Without this deep copy, some pointers are shared between several servers, leading to UAF, double-free or such bugs. This patch relies on following commits: * `b32cb9b51` REORG: server: Export srv_settings_cpy() function * `0b365e3cb` MINOR: server: Constify source server to copy its settings This patch should fix the issue #1804. It must be backported as far as 2.0.	2022-08-03 11:44:34 +02:00
Christopher Faulet	b32cb9b515	REORG: server: Export srv_settings_cpy() function This function will be used to init a proxy with settings of the default proxy. It is mandatory to fix a bug. To do so, it must be exposed.	2022-08-03 11:28:52 +02:00
Christopher Faulet	0b365e3cb5	MINOR: server: Constify source server to copy its settings The source server used to initialize a new server, in srv_settings_cpy() and sub-functions, is now a constant. This patch is mandatory to fix a bug.	2022-08-03 11:28:23 +02:00
Christopher Faulet	bc6b23813f	BUG/MINOR: backend: Don't increment conn_retries counter too early The connection retry counter is incremented too early when a connection fails. In SC_ST_CER state, errors handling must be performed before incrementing the counter. Otherwise, we may consider the max connection attempt is reached while a last one is in fact possible. This patch must be backported to 2.6.	2022-08-03 11:16:35 +02:00
Christopher Faulet	14a60d420a	BUG/MEDIUM: dns: Properly initialize new DNS session When a new DNS session is created, all its fields are not properly initialized. For instance, "tx_msg_offset" can have any value after the allocation. So, to fix the bug, pool_zalloc() is now used to allocate new DNS session. This patch should fix the issue #1781. It must be backported as far as 2.4.	2022-08-03 10:30:07 +02:00
Christopher Faulet	642170a653	BUG/MINOR: peers: Use right channel flag to consider the peer as connected When a peer open a new connection to another peer, it is considered as connected when the hello message is sent. To do so, the peer applet was relying on CF_WRITE_PARTIAL channel flag. However it is not the right flag to use. This one is a transient flag. Depending on the scheduling, this flag may be removed by the stream before the peer has a chance to see it. Instead, CF_WROTE_DATA flag must be checked. This patch is related to the issue #1799. It must be backported as far as 2.0.	2022-08-03 09:56:38 +02:00
Christopher Faulet	160fff665e	BUG/MEDIUM: peers: limit reconnect attempts of the old process on reload When peers are configured and HAProxy is reloaded or restarted, a synchronization is performed between the old process and the new one. To do so, the old process connects on the new one. If the synchronization fails, it retries. However, there is no delay and reconnect attempts are not bounded. Thus, it may loop for a while, consuming all the CPU. Of course, it is unexpected, but it is possible. For instance, if the local peer is misconfigured, an infinite loop can be observed if the connection succeeds but not the synchronization. This prevents the old process to exit, except if "hard-stop-after" option is set. To fix the bug, the reconnect is delayed. The local peer already has a expiration date to delay the reconnects. But it was not used on stopping mode. So we use it not. Thanks to the previous fix, the reconnect timeout is shorter in this case (500ms against 5s on running mode). In addition, we also use the peers resync expiration date to not infinitely retries. It is accurate because the new process, on its side, use this timeout to switch from a local resync to a remote resync. This patch depends on "MINOR: peers: Use a dedicated reconnect timeout when stopping the local peer". It fixes the issue #1799. It should be backported as far as 2.0.	2022-08-03 09:56:38 +02:00
Christopher Faulet	ab4b094055	MINOR: peers: Use a dedicated reconnect timeout when stopping the local peer When a process is stopped or reload, a dedicated reconnect timeout is now used. For now, this timeout is not used because the current code retries immediately to reconnect to perform the local synchronization with the new local peer, if any. This patch is required to fix the issue #1799. It should be backported as far as 2.0 with next fixes.	2022-08-03 09:56:38 +02:00
Christopher Faulet	1b6fa7f5ea	MINOR: peers: Add a warning about incompatible SSL config for the local peer In peers section, it is possible to enable SSL for the local peer. In this case, the bind line and the server line should both be configured. A "default-server" directive may also be used to configure the SSL on the server side. However there is no test to be sure the SSL is enabled on both sides. It is an problem because the local resync performed during a reload will be impossible and it is probably not the expected behavior. So, it is now checked during the configuration validation. A warning message is displayed if the SSL is not properly configured for the local peer. This patch is related to issue #1799. It should probably be backported to 2.6.	2022-08-03 09:56:38 +02:00
Amaury Denoyelle	bd6ec1bf84	MEDIUM: mux-quic: implement http-keep-alive timeout Complete QUIC MUX timeout refresh function by using http-keep-alive timeout. It is used when the connection is idle after having handle at least one request. To implement this a new member <idle_start> has been defined in qcc structure. This is used as timestamp for when the connection became idle and is used as base time for http keep-alive timeout	2022-08-01 15:00:13 +02:00
Amaury Denoyelle	c603de4d84	MINOR: mux-quic: count in-progress requests Add a new qcc member named <nb_hreq>. Its purpose is close to <nb_sc> which represents the number of attached stream connectors. Both are incremented inside qc_attach_sc(). The difference is on the decrement operation. While <nb_cs> is decremented on sedesc detach callback, <nb_hreq> is decremented when the qcs is locally closed. In most cases, <nb_hreq> will be decremented before <nb_cs>. However, it will be the reverse if a stream must be kept alive after detach callback. The main purpose of this field is to implement http-keep-alive timeout. Both <nb_sc> and <nb_hreq> must be null to activate the http-keep-alive timeout.	2022-08-01 14:58:41 +02:00
Amaury Denoyelle	5fc05d17ad	MEDIUM: mux-quic: adjust timeout refresh Implement a new internal function qcc_refresh_timeout(). Its role will be to reset QUIC MUX timeout depending if there is requests in progress or not. qcc_update_timeout() does not set a timeout if there is still attached streams as in this case the upper layer is responsible to manage it. Else it will activate the timeout depending on the connection current status. Timeout is refreshed on several locations : on stream detach and in I/O handler and wake callback. For the moment, only the default timeout is used (client or server). The function may be expanded in the future to support more specific ones : * http-keep-alive if connection is idle * http-request when waiting for incomplete HTTP requests * client/server-fin for graceful shutdown	2022-08-01 14:58:36 +02:00
Amaury Denoyelle	b6309456d0	MINOR: mux-quic: use timeout server for backend conns Use timeout server in qcc_init() as default timeout for backend connections. No impact for the moment as QUIC backend support is not implemented.	2022-08-01 14:23:21 +02:00
Amaury Denoyelle	07bf8f4d86	MINOR: mux-quic: save proxy instance into qcc Store a reference to proxy in the qcc structure. This will be useful to access to proxy members outside of qcc_init(). Most notably, this change is required to implement timeout refreshing by using the various timeouts configured at the proxy level.	2022-08-01 14:23:21 +02:00
Amaury Denoyelle	09ec3e09bd	BUG/MINOR: mux-quic: do not free conn if attached streams Ensure via qcc_is_dead() that a connection is not released instance until all of qcs streams are detached by the upper layer, even if an error has been reported or the timeout has fired. On the other side, as qc_detach() always check the connection status, this should ensure that we do not keep a connection if not necessary. Without this patch, a qcc instance may be freed with some of its qcs streams not detached. This is an incorrect behavior and will lead to a BUG_ON fault. Note however that no occurence of this bug has been produced currently. This patch is mainly a safety against future occurences. This should be backported up to 2.6.	2022-08-01 14:23:19 +02:00
Amaury Denoyelle	4ea5090f55	CLEANUP: mux-quic: remove useless app_ops is_active callback Timeout in QUIC MUX has evolved from the simple first implementation. At the beginning, a connection was considered dead unless bidirectional streams were opened. This was abstracted through an app callback is_active(). Now this paradigm has been reversed and a connection is considered alive by default, unless an error has been reported or a timeout has already been fired. The callback is_active() is thus not used anymore and can be safely removed to simplify qcc_is_dead(). This commit should be backported to 2.6.	2022-08-01 14:13:51 +02:00
Amaury Denoyelle	d3973853c2	BUG/MINOR: mux-quic: prevent crash if conn released during IO callback A qcc instance may be freed in the middle of qc_io_cb() if all streams were purged. This will lead to a crash as qcc instance is reused after this step. Jump directly to the end of the function to avoid this. Note that this bug has not been triggered for the moment. This is a safety fix to prevent it. This must be backported up to 2.6.	2022-08-01 14:13:51 +02:00
Willy Tarreau	51d38a26fe	BUG/MEDIUM: pattern: only visit equivalent nodes when skipping versions Miroslav reported in issue #1802 a problem that affects atomic map/acl updates. During an update, incorrect versions are properly skipped, but in order to do so, we rely on ebmb_next() instead of ebmb_next_dup(). This means that if a new matching entry is in the process of being added and is the first one to succeed in the lookup, we'll skip it due to its version and use the next entry regardless of its value provided that it has the correct version. For IP addresses and string prefixes it's particularly visible because a lookup may match a new longer prefix that's not yet committed (e.g. 11.0.0.1 would match 11/8 when 10/7 was the only committed one), and skipping it could end up on 12/8 for example. As soon as a commit for the last version happens, the issue disappears. This problem only affects tree-based matches: the "str", "ip", and "beg" matches. Here we replace the ebmb_next() values with ebmb_next_dup() for exact string matches, and with ebmb_lookup_shorter() for longest matches, which will first visit duplicates, then look for shorter prefixes. This relies on previous commit: MINOR: ebtree: add ebmb_lookup_shorter() to pursue lookups Both need to be backported to 2.4, where the generation ID was added. Note that nowadays a simpler and more efficient approach might be employed, by having a single version in the current tree, and a list of trees per version. Manipulations would look up the tree version and work (and lock) only in the relevant trees, while normal operations would be performed on the current tree only. Committing would just be a matter of swapping tree roots and deleting old trees contents.	2022-08-01 11:59:46 +02:00
Willy Tarreau	0dc9e6dca2	DEBUG: tools: provide a tree dump function for ebmbtrees as well It's convenient for debugging IP trees. However we're not dumping the full keys, for the sake of simplicity, only the 4 first bytes are dumped as a u32 hex value. In practice this is sufficient for debugging. As a reminder since it seems difficult to recover the command each time it's needed, the output is converted to an image using dot from Graphviz: dot -o a.png -Tpng dump.txt	2022-08-01 11:59:15 +02:00
Willy Tarreau	87aff021db	MINOR: thread: provide an alternative to pthread's rwlock Since version 1.1.0, OpenSSL's libcrypto ignores the provided locking mechanism and uses pthread's rwlocks instead. The problem is that for some code paths (e.g. async engines) this results in a huge amount of syscalls on systems facing a bit of contention, to the point where more than 80% of the CPU can be spent in the system dealing with spinlocks just for futex_wake(). This patch provides an alternative by redefining the relevant pthread rwlocks from the low-overhead version of the progressive rw locks. This way there will be no more syscalls in case of contention, and CPU will be burnt in userland. Doing this saves massive amounts of CPU, where the locks only take 12-15% vs 80% before, which allows SSL to work much faster on large thread counts (e.g. 24 or more). The tryrdlock and trywrlock variants have been implemented using a CAS since their goal is only to succeed on no contention and never to wait. The pthread_rwlock API is complete except that the timed versions of the rdlock and wrlock do not wait and simply fall back to trylock versions. Since the gains have only been observed with async engines for now, this option remains disabled by default. It can be enabled at build time using USE_PTHREAD_EMULATION=1.	2022-07-30 10:17:22 +02:00
Willy Tarreau	ddab05b98a	BUG/MEDIUM: queue/threads: limit the number of entries dequeued at once When testing strong queue contention on a 48-thread machine, some crashes would frequently happen due to process_srv_queue() never leaving and processing pending requests forever. A dump showed more than 500000 loops at once. The problem is that other threads find it working so they don't do anything and are free to process their pending requests. Because of this, the dequeuing thread can be kept busy forever and does not process its own requests anymore (fortunately the watchdog stops it). This patch adds a limit to the number of rounds, it limits it to maxpollevents, which is reasonable because it's also an indicator of latency and batches size. However there's a catch. If all requests are finished when the thread ends the loop, there might not be enough anymore to restart processing the queue. Thus we tolerate to re-enter the loop to process one request at a time when it doesn't have any anymore. This way we're leaving more room for another thread to take on this task, and we're sure to eventually end this loop. Doing this has also improved the overall dequeuing performance by ~20% in highly contended situations with 48 threads. It should be backported at least to 2.4, maybe even 2.2 since issues were faced in the past on machines having many cores.	2022-07-30 10:00:59 +02:00
Frédéric Lécaille	dc07751ed7	MINOR: quic: Send packets as much as possible from qc_send_app_pkts() Add a loop into this function to send more packets from this function which is called by the mux. It is broken when we could not prepare packet with qc_prep_app_pkts() due to missing available room in the buffer used to send packets. This improves the throughput. Must be backported to 2.6.	2022-07-29 17:32:05 +02:00
Frédéric Lécaille	843399fd45	BUG/MAJOR: quic: Useless resource intensive loop qc_ackrng_pkts() This usless loop should have been removed a long time ago. As it is CPU resource intensive, it could trigger the watchdog. Must be backported to 2.6.	2022-07-29 17:32:05 +02:00
Frédéric Lécaille	dc591cd6cb	MINOR: quic: Stop looking for packet loss asap As the TX packets are ordered by their packet number and always sent in the same order. their TX timestamps are inspected from the older to the newer values when we look for the packet loss. So we can stop this search as soon as we found the first packet which has not been lost. Must be backported to 2.6	2022-07-29 17:32:05 +02:00
Frédéric Lécaille	d2e104ff78	BUG/MINOR: quic: loss time limit variable computed but not used <loss_time_limit> is the loss time limit computed from <time_sent> packet transmission timestamps in qc_packet_loss_lookup() to identify the packets which have been lost. This latter timestamp variable was used in place of <loss_time_limit> to distinguish such packets from others (still in fly packets). Must be backported to 2.6	2022-07-29 17:32:05 +02:00
Frédéric Lécaille	43910a9450	MINOR: quic: New "quic-cc-algo" bind keyword As it could be interesting to be able to choose the QUIC control congestion algorithm to be used by listener, add "quic-cc-algo" new keyword to do so. Update the documentation consequently. Must be backported to 2.6.	2022-07-29 17:32:05 +02:00
Frédéric Lécaille	1c9c2f6c02	MEDIUM: quic: Cubic congestion control algorithm implementation Cubic is the congestion control algorithm used by default by the Linux kernel since 2.6.15 version. This algorithm is supposed to achieve good scalability and fairness between flows using the same network path, it should also be used by QUIC by default. This patch implements this algorithm and select it as default algorithm for the congestion control. Must be backported to 2.6.	2022-07-29 17:32:05 +02:00
Frédéric Lécaille	c591459d11	MINOR: quic: Congestion control architecture refactoring Ease the integration of new congestion control algorithm to come. Move the congestion controller state to a private array of uint32_t to stop using a union. We do not want to continue using such long paths cc->algo_state.<algo>.<var> to modify the internal state variable for each algorithm. Must be backported to 2.6	2022-07-29 17:32:05 +02:00
Amaury Denoyelle	72a78e8290	BUG/MEDIUM: mux-quic: fix missing EOI flag to prevent streams leaks On H3 DATA frame transfer from the client, some streams are not properly closed by the upper layer, despite all transfer operation completed. Data integrity is not impacted but this will prevent the stream timeout to fire and thus keep the owner session opened. In most cases, sessions are closed on QUIC idle timeout, but it may stay forever if a client emits PING frames at a regular interval to maintain it. This bug is caused by a missing EOI stream desc flag on certain condition in qc_rcv_buf(). To be triggered, we have to use the optimization when conn-stream buffer is empty and can be swapped with qcs buffer. The problem is that it will skip the function body for default copy but also a condition to check if EOI must be set. Thus this bug does not happens for every H3 post requets : it requires that conn-stream buffer is empty on last qc_rcv_buf() invocation. This was reproduced more frequently when using ngtcp2 client with one or multiple streams : $ ngtcp2-client -m POST -d ~/infra/html/10K 127.0.0.1 20443 \ http://127.0.0.1:20443/post This may fix at least partially github issue #1801. This must be backported up to 2.6.	2022-07-29 16:01:21 +02:00
William Lallemand	b5d062dff1	MINOR: cli: warning on _getsocks when socket were closed The previous attempt was reverted because it would emit a warning when the sockets are still in the process when a reload failed, so this was an expected 2nd try. This warning however, will be displayed if a new process successfully get the previous sockets AND the sendable number of sockets is 0. This way the user will be warned if he tried to get the sockets fromt the wrong process.	2022-07-28 15:49:43 +02:00
William Lallemand	9c821e615e	Revert "MINOR: cli: emit a warning when _getsocks was used more than once" This reverts commit `519cd2021b`. This was reverted because it's still useful to have access to _getsosks when the previous reload failed.	2022-07-27 13:55:54 +02:00
William Lallemand	14b98ef1bd	BUG/MINOR: mworker: PROC_O_LEAVING used but not updated Since commit `2be557f` ("MEDIUM: mworker: seamless reload use the internal sockpair"), we are using the PROC_O_LEAVING flag to determine which sockpair worker will be used with -x during the next reload. However in mworker_reexec(), the PROC_O_LEAVING flag is not updated, it is only updated at startup in mworker_env_to_proc_list(). This could be a problem when a remaining process is still in the list, it could be selected as the current worker, and its socket will be used even if _getsocks doesn't work anymore on it. (bug #1803) This patch fixes the issue by updating the PROC_O_LEAVING flag in mworker_proc_list_to_env() just before using it in mworker_reexec() Must be backported to 2.6.	2022-07-27 12:13:56 +02:00
William Lallemand	519cd2021b	MINOR: cli: emit a warning when _getsocks was used more than once The _getsocks CLI command can be used only once, after that the sockets are not available anymore. Emit a warning when the command was already used once.	2022-07-27 11:48:54 +02:00
Willy Tarreau	b983145837	BUG/MINOR: fd: always remove late updates when freeing fd_updt[] Christopher found that since commit `8e2c0fa8e` ("MINOR: fd: delete unused updates on close()") we may crash in a late stop due to an fd_delete() in the main thread performed after all threads have deleted the fd_updt[] array. Prior to that commit that didn't happen because we didn't touch the updates on this path, but now it may happen. We don't care about these ones anyway since the poller is stopped, so let's just wipe them by resetting their counter before freeing the array. No backport is needed as this is only 2.7.	2022-07-26 19:06:17 +02:00
William Lallemand	c31577f32e	MEDIUM: resolvers: continue startup if network is unavailable When haproxy starts with a resolver section, and there is a default one since 2.6 which use /etc/resolv.conf, it tries to do a connect() with the UDP socket in order to check if the routes of the system allows to reach the server. This check is too much restrictive as it won't prevent any runtime failure. Relax the check by making it a warning instead of a fatal alert. This must be backported in 2.6.	2022-07-26 10:59:14 +02:00
Christopher Faulet	244331f6e7	Revert "BUG/MINOR: peers: set the proxy's name to the peers section name" This reverts commit `356866acce`. It seems that an undocumented expectation of peers is based on the peers proxy name to determine if the local peer is fully configured or not. Thus because of the commit above, we are no longer able to detect incomplete peers sections. On side effect of this bug is a segfault when HAProxy is stopped/reloaded if we try to perform a local resync on a mis-configured local peer. So waiting for a better solution, the patch is reverted. This patch must be backported as far as 2.5.	2022-07-25 16:17:04 +02:00
William Lallemand	708949da49	MINOR: sockpair: move send_fd_uxst() error message in caller Move the ha_alert() in send_fd_uxst() in the callers and add the FD numbers in the message.	2022-07-25 16:11:11 +02:00
William Lallemand	f67e8fb92c	BUG/MINOR: sockpair: wrong return value for fd_send_uxst() The fd_send_uxst() function which is used to send a socket over the socketpair returns 1 upon error instead of -1, which means the error case of the sendmsg() is never catched correctly. Must be backported as far as 1.9.	2022-07-25 16:10:58 +02:00
Willy Tarreau	6983426354	BUG/MAJOR: poller: drop FD's tgid when masks don't match A bug was introduced in 2.7-dev2 by commit `1f947cb39` ("MAJOR: poller: only touch/inspect the update_mask under tgid protection"): once the FD's tgid is held, we would forget to drop it in case the update mask doesn't match, resulting in random watchdog panics of older processes on successive reloads. This should fix issue #1798. Thanks to Christian for the report and to Christopher for the reproducer. No backport is needed.	2022-07-25 15:47:15 +02:00
Willy Tarreau	53bfac8c63	BUG/MEDIUM: master: force the thread count earlier Christopher bisected that recent commit `d0b73bca71` ("MEDIUM: listener: switch bind_thread from global to group-local") broke the master socket in that only the first out of the Nth initial connections would work, where N is the number of threads, after which they all work. The cause is that the master socket was bound to multiple threads, despite global.nbthread being 1 there, so the incoming connection load balancing would try to send incoming connections to non-existing threads, however the bind_thread mask would nonetheless include multiple threads. What happened is that in 1.9 we forced "nbthread" to 1 in the master's poll loop with commit `b3f2be338b` ("MEDIUM: mworker: use the haproxy poll loop"). In 2.0, nbthread detection was enabled by default in commit `149ab779cc` ("MAJOR: threads: enable one thread per CPU by default"). From this point on, the operation above is unsafe because everything during startup is performed with nbthread corresponding to the default value, then it changes to one when starting the polling loop. But by then we weren't using the wait mode except for reload errors, so even if it would have happened nobody would have noticed. In 2.5 with commit `fab0fdce9` ("MEDIUM: mworker: reexec in waitpid mode after successful loading") we started to rexecute all the time, not just for errors, so as to release precious resources and to possibly spot bugs that were rarely exposed in this mode. By then the incoming connection LB was enforcing all_threads_mask on the listener's thread mask so that the incorrect value was being corrected while using it. Finally in 2.7 commit `d0b73bca71` ("MEDIUM: listener: switch bind_thread from global to group-local") replaces the all_threads_mask there with the listener's bind_thread, but that one was never adjusted by the starting master, whose thread group was filled to N threads by the automatic detection during early setup. The best approach here is to set nbthread to 1 very early in init() when we're in the master in wait mode, so that we don't try to guess the best value and don't end up with incorrect bindings anymore. This patch does this and also sets nbtgroups to 1 in preparation for a possible future where this will also be automatically calculated. There is no need to backport this patch since no other versions were affected, but if it were to be discovered that the incorrect bind mask on some of the master's FDs could be responsible for any trouble in older versions, then the backport should be safe (provided that nbtgroups is dropped of course).	2022-07-22 17:51:53 +02:00
Christopher Faulet	38c53944cb	BUG/MINOR: backend: Fallback on RR algo if balance on source is impossible If the loadbalancing is performed on the source IP address, an internal error was returned on error. So for an applet on the client side (for instance an SPOE applet) or for a client connected to a unix socket, an internal error is returned. However, when other LB algos fail, a fallback on round-robin is performed. There is no reson to not do the same here. This patch should fix the issue #1797. It must be backported to all supported versions.	2022-07-22 17:07:34 +02:00
Christopher Faulet	ca67992979	BUG/MEDIUM: stconn: Only reset connect expiration when processing backend side Since commit `ae024ced0` ("MEDIUM: stream-int/stream: Use connect expiration instead of SI expiration"), the connect expiration date is per-stream. So there is only one expiration date instead of one per side, front and back. So when a stream-connector is processed, we must test if it is a frontend or a backend stconn before updating the connect expiration date. Indeed, the frontend stconn must not reset the connect expiration date. This bug may have several side effect. One known bug is about peer sessions blocked because the frontend peer applet is in ST_CLO state and its backend connection is in ST_TAR state but without connect expiration date. This patch should fix the issue #1791 and #1792. It must be backported to 2.6.	2022-07-21 14:50:14 +02:00
Willy Tarreau	41afd9084e	BUILD: add detection for unsupported compiler models As reported in github issue #1765, some people get trapped into building haproxy and companion libraries on Windows using a compiler following the LLP64 model. This has no chance to work, and definitely causes nasty bugs everywhere when pointers are passed as longs. Let's save them time and detect this at boot time. The message and detection was factored with the existing one for -fwrapv since we need the same info and actions. This should be backported to all recent supported versions (the ones that are likely to be tried on such platforms when people don't know).	2022-07-21 09:58:20 +02:00
William Lallemand	d4835a9680	BUG/MEDIUM: mworker: proc_self incorrectly set crashes upon reload When updating from 2.4 to 2.6, the child->reloads++ instruction changed place, resulting in a former worker from the 2.4 process, still identified as a current worker once in 2.6, because its reload counter is still 0. Unfortunately this counter is used to chose the mworker_proc structure that will be used for the new worker. What happens next, is that the mworker_proc structure of the previous process is selected, and this one has ipc_fd[1] set to -1, because this structure was supposed to be in the master. The process then forks, and mworker_sockpair_register_per_thread() tries to register ipc_fd[1] which is set to -1, instead of the fd of the new socketpair. This patch fixes the issue by checking if child->pid is equal to -1 when selecting proc_self. This way we could be sure it wasn't a previous process. Should fix issue #1785. This must be backported as far as 2.4 to fix the issue related to the reload computation difference. However backporting it in every stable branch will enforce the reload process.	2022-07-21 00:52:43 +02:00
Frédéric Lécaille	a18c3339c8	BUG/MAJOR: mux_quic: fix invalid PROTOCOL_VIOLATION on POST data overlap Stream data reception is incorrect when dealing with a partially new offset with some data already consumed out of the RX buffer. In this case, data length is adjusted but not the data buffer. In most cases, ncb_add() operation will be rejected as already stored data does not correspond with the new inserted offset. This will result in an invalid CONNECTION_CLOSE with PROTOCOL_VIOLATION. To fix this, buffer pointer is advanced while the length is reduced. This can be reproduced with a POST request and patching haproxy to call qcc_recv() multiple times by copying a quic_stream frame with different offsets. Must be backported to 2.6.	2022-07-20 15:34:58 +02:00
William Lallemand	bac3a82a50	BUG/MINOR: mworker/cli: relative pid prefix not validated anymore Since `e8422bf` ("MEDIUM: global: remove the relative_pid from global and mworker"), the relative pid prefix is not tested anymore on the master CLI. Which means any value will fall into the "1" process. Since we removed the nbproc, only the "1" and the "0" (master) value are correct, any other value should return an error. Fix issue #1793. This must be backported as far as 2.5.	2022-07-20 14:43:47 +02:00
William Lallemand	0f17ab2fdd	MINOR: ssl: enhance ca-file error emitting Enhance the errors and warnings when trying to load a ca-file with ssl_store_load_locations_file(). Add errors from ERR_get_error() and strerror to give more information to the user.	2022-07-19 19:13:08 +02:00
William Lallemand	3b8bafd4a7	MINOR: init: load OpenSSL error strings Load OpenSSL Error strings in order to be able to output reason strings. This is mandatory to be able to use ERR_reason_error_string().	2022-07-19 19:13:08 +02:00
Willy Tarreau	c1640f79fe	BUG/MEDIUM: fd/threads: fix incorrect thread selection in wakeup broadcast In commit `cfdd20a0b` ("MEDIUM: fd: support broadcasting updates for foreign groups in updt_fd_polling") we decided to pick a random thread number among a set of candidates for a wakeup in case we need an instant change. But the thread count range was wrong (MAX_THREADS) instead of tg->count, resulting in random crashes when thread groups are > 1 and MAX_THREADS > 64. No backport is needed, this was introduced in 2.7-dev2.	2022-07-19 16:01:04 +02:00
Christopher Faulet	f7ebe584d7	BUILD: debug: Add braces to if statement calling only CHECK_IF() In src/ev_epoll.c, a CHECK_IF() is guarded by an if statement. So, when the macro is empty, GCC (at least 11.3.1) is not happy because there is an if statement with an empty body without braces... It is handled by "-Wempty-body" option. So, braces are added and GCC is now happy. No backport needed.	2022-07-19 12:11:04 +02:00
Amaury Denoyelle	0933c7b3c8	BUG/MINOR: quic: do not send CONNECTION_CLOSE_APP in initial/handshake As specified by RFC 9000, it is forbidden to send a CONNECTION_CLOSE of type 0x1d (CONNECTION_CLOSE_APP) in an Initial or Handshake packet. It must be converted to type 0x1c (CONNECTION_CLOSE) with APPLICATION_ERROR code. CONNECTION_CLOSE_APP are generated by QUIC MUX interaction. Thus, special care must be taken when dealing with a 0-RTT packet, as this is the only case where the MUX can be instantiated and quic-conn still on the Initial or Handshake encryption level. To enforce RFC 9000, xprt build packet function is now responsible to translate a CONNECTION_CLOSE_APP if still on Initial/Handshake encryption. This process is done in a dedicated function named qc_build_cc_frm(). Without this patch, BUG_ON() statement in qc_build_frm() will be triggered when building a CONNECTION_CLOSE_APP frame on Initial or Handshake level. This is because QUIC_FT_CONNECTION_CLOSE_APP frame builder mask does not allow these encryption levels, as opposed to QUIC_FT_CONNECTION_CLOSE builder. This crash was reproduced by modifying the H3 layer to force emission of a CONNECTION_CLOSE_APP on first frame of a 0-RTT session. Note however that CONNECTION_CLOSE emission during Handshake is a complicated process for the server. For the moment, this is still incomplete on haproxy side. RFC 9000 requires to emit it multiple times in several packets under different encryption levels, depending on what we know about the client encryption context. This patch should be backported up to 2.6.	2022-07-19 11:19:50 +02:00
William Lallemand	4348232231	BUG/MINOR: ssl: allow duplicate certificates in ca-file directories It looks like OpenSSL 1.0.2 returns an error when trying to insert a certificate whis is already present in a X509_STORE. This patch simply ignores the X509_R_CERT_ALREADY_IN_HASH_TABLE error if emitted. Should fix part of issue #1780. Must be backported in 2.6.	2022-07-18 18:49:27 +02:00
William Lallemand	3bda80789c	BUG/MINOR: resolvers: shut off the warning for the default resolvers When the resolv.conf file is empty or there is no resolv.conf file, an empty resolvers will be created, which emits a warning during the postparsing step. This patch fixes the problem by freeing the resolvers section if the parsing failed or if the nameserver list is empty. Must be backported in 2.6, the previous patch which introduces resolvers_destroy() is also required.	2022-07-18 14:39:36 +02:00
William Lallemand	e606c84fee	MINOR: resolvers: resolvers_destroy() deinit and free a resolver Split the resolvers_deinit() function into resolvers_destroy() and resolvers_deinit() in order to be able to free a unique resolvers section.	2022-07-18 14:39:36 +02:00
Willy Tarreau	5b3cd9561b	BUG/MEDIUM: tools: avoid calling dlsym() in static builds (try 2) The first approach in commit `288dc1d8e` ("BUG/MEDIUM: tools: avoid calling dlsym() in static builds") relied on dlopen() but on certain configs (at least gcc-4.8+ld-2.27+glibc-2.17) it used to catch situations where it ought not fail. Let's have a second try on this using dladdr() instead. The variable was renamed "build_is_static" as it's exactly what's being detected there. We could even take it for reporting in -vv though that doesn't seem very useful. At least the variable was made global to ease inspection via the debugger, or in case it's useful later. Now it properly detects a static build even with gcc-4.4+glibc-2.11.1 and doesn't crash anymore.	2022-07-18 14:03:54 +02:00
Willy Tarreau	288dc1d8ee	BUG/MEDIUM: tools: avoid calling dlsym() in static builds Since 2.4 with commit `64192392c` ("MINOR: tools: add functions to retrieve the address of a symbol"), we can resolve symbols. However some old glibc crash in dlsym() when the program is statically built. Fortunately even on these old libs we can detect lack of support by calling dlopen(NULL). Normally it returns a handle to the current program, but on a static build it returns NULL. This is sufficient to refrain from calling dlsym() (which will be of very limited use anyway), so we check this once at boot and use the result when needed. This may be backported to 2.4. On stable versions, be careful to place the init code inside an if/endif guard that checks for DL support.	2022-07-16 13:49:34 +02:00
Willy Tarreau	c6b596dcce	CLEANUP: threads: remove the now unused all_threads_mask and tid_bit Since these are not used anymore, let's now remove them. Given the number of places where we're using ti->ldit_bit, maybe an equivalent might be useful though.	2022-07-15 20:25:41 +02:00
Willy Tarreau	cfdd20a0b2	MEDIUM: fd: support broadcasting updates for foreign groups in updt_fd_polling We're still facing the situation where it's impossible to update an FD for a foreign group. That's of particular concern when disabling/enabling listeners (e.g. pause/resume on signals) since we don't decide which thread gets the signal and it needs to process all listeners at once. Fortunately, not that much is unprotected in FDs. This patch adds a test for tgid's equality in updt_fd_polling() so that if a change is applied for a foreing group, then it's detected and taken care of separately. The method consists in forcing the update on all bound threads in this group, adding it to the group's update_list, and sending a wake-up as would be done for a remote thread in the local group, except that this is done by grabbing a reference to the FD's tgid. Thanks to this, SIGTTOU/SIGTTIN now work for nbtgroups > 1 (after that was temporarily broken by "MEDIUM: fd/poller: make the update-list per-group").	2022-07-15 20:25:41 +02:00
Willy Tarreau	1f947cb39e	MAJOR: poller: only touch/inspect the update_mask under tgid protection With thread groups and group-local masks, the update_mask cannot be touched nor even checked if it may change below us. In order to avoid this, we have to grab a reference to the FD's tgid before checking the update mask. The operations are cheap enough so that we don't notice it in performance tests. This is expected because the risk of meeting a reassigned FD during an update remains very low. It's worth noting that the tgid cannot be trusted during startup nor during soft-stop since that may come from anywhere at the moment. Since soft-stop runs under thread isolation we use that hint to decide whether or not to check that the FD's tgid matches the current one. The modification is applied to the 3 thread-aware pollers, i.e. epoll, kqueue, and evports. Also one poll_drop counter was missing for shared updates, though it might be hard to trigger it. With this change applied, thread groups are usable in benchmarks.	2022-07-15 20:16:30 +02:00
Willy Tarreau	d95f18fa39	MAJOR: pollers: rely on fd_reregister_all() at boot time The poller-specific thread init code now uses that new function to safely register boot events. This ensures that we don't register an event for another group and that we properly deal with parallel thread startup. It's only done for thread-aware pollers, there's no point in using that in poll/select though that should work as well.	2022-07-15 20:16:30 +02:00
Willy Tarreau	9baff4ffd9	MEDIUM: fd: support stopping FDs during starting There's a nasty case during boot, which is the master process. It stops all listeners from the main thread, and as such we're seeing calls to fd_delete() from a thread that doesn't match the FD's mask, but more importantly from a group that doesn't match either. Fortunately this happens in a process that doesn't see the threads creation, so the FDs are left intact in the table and we can overwrite the tgid there. The approach is ugly, it probably shows that we should use a dummy value for the tgid during boot, that would be replaced once the FDs migrate to their target, but we also need a way to make sure not to miss them. Also that doesn't solve the possibility of closing a listener at run time from the wrong thread group.	2022-07-15 20:16:30 +02:00
Willy Tarreau	88c4c14050	MINOR: fd: add fd_reregister_all() to deal with boot-time FDs At boot the pollers are allocated for each thread and they need to reprogram updates for all FDs they will manage. This code is not trivial, especially when trying to respect thread groups, so we'd rather avoid duplicating it. Let's centralize this into fd.c with this function. It avoids closed FDs, those whose thread mask doesn't match the requested one or whose thread group doesn't match the requested one, and performs the update if required under thread-group protection.	2022-07-15 20:16:30 +02:00
Willy Tarreau	d0b73bca71	MEDIUM: listener: switch bind_thread from global to group-local It requires to both adapt the parser and change the algorithm to redispatch incoming traffic so that local threads IDs may always be used. The internal structures now only reference thread group IDs and group-local masks which are compatible with those now used by the FD layer and the rest of the code.	2022-07-15 20:16:30 +02:00
Willy Tarreau	6018c02c36	MEDIUM: thread: change thread_resolve_group_mask() to return group-local values It used to turn group+local to global but now we're doing the exact opposite as we want to stick to group-local masks. This means that "thread 3-4" might very well emit what "thread 2/1-2" used to emit till now for 2 groups and 4 threads. This is needed because we'll have to support group-local thread masks in receivers. However the rest of the code (receivers) is not ready yet for this, so using this code with more than one thread group will definitely break some bindings.	2022-07-15 20:16:30 +02:00
Willy Tarreau	0b51eab764	MEDIUM: fd: quit fd_update_events() when FD is closed The IOCB might have closed the FD itself, so it's not an error to have fd.tgid==0 or anything else, nor to have a null running_mask. In fact there are different conditions under which we can leave the IOCB, all of them have been enumerated in the code's comments (namely FD still valid and used, hence has running bit, FD closed but not yet reassigned thus running==0, FD closed and reassigned, hence different tgid and running becomes irrelevant, just like all other masks). For this reason we have no other solution but to try to grab the tgid on return before checking the other bits. In practice it doesn't represent a big cost, because if the FD was closed and reassigned, it's instantly detected and the bit is immediately released without blocking other threads, and if the FD wasn't closed this doesn't prevent it from being migrated to another thread. In the worst case a close by another thread after a migration will be postponed till the moment the running bit is cleared, which is the same as before.	2022-07-15 20:16:30 +02:00
Willy Tarreau	ddedc16624	MEDIUM: fd: make fd_insert/fd_delete atomically update fd.tgid These functions need to set/reset the FD's tgid but when they're called there may still be wakeups on other threads that discover late updates and have to touch the tgid at the same time. As such, it is not possible to just read/write the tgid there. It must only be done using operations that are compatible with what other threads may be doing. As we're using inc/dec on the refcount, it's safe to AND the area to zero the lower part when resetting the value. However, in order to set the value, there's no other choice but fd_claim_tgid() which will assign it only if possible (via a CAS). This is convenient in the end because it protects the FD's masks from being modified by late threads, so while we hold this refcount we can safely reset the thread_mask and a few other elements. A debug test for non-null masks was added to fd_insert() as it must not be possible to face this situation thanks to the protection offered by the tgid.	2022-07-15 20:16:30 +02:00
Willy Tarreau	27a3245599	MEDIUM: fd: make fd_insert() take local thread masks fd_insert() was already given a thread group ID and a global thread mask. Now we're changing the few callers to take the group-local thread mask instead. It's passed directly into the FD's thread mask. Just like for previous commit, it must not change anything when a single group is configured.	2022-07-15 20:16:30 +02:00
Willy Tarreau	3638d174e5	MEDIUM: fd: make thread_mask now represent group-local IDs With the change that was started on other masks, the thread mask was still not fully converted, sometimes being used as a global mask and sometimes as a local one. This finishes the code modifications so that the mask is always considered as a group-local mask. This doesn't change anything as long as there's a single group, but is necessary for groups 2 and above since it's used against running_mask and so on.	2022-07-15 20:16:30 +02:00
Willy Tarreau	d6e1987612	MINOR: fd: make fd_clr_running() return the previous value instead It's an AND so it destroys information and due to this there's a call place where we have to perform two reads to know the previous value then to change it. With a fetch-and-and instead, in a single operation we can know if the bit was previously present, which is more efficient.	2022-07-15 20:16:30 +02:00
Willy Tarreau	a707d02657	MEDIUM: fd/poller: turn running_mask to group-local IDs From now on, the FD's running_mask only refers to local thread IDs. However, there remains a limitation, in updt_fd_polling(), we temporarily have to check and set shared FDs against .thread_mask, which still contains global ones. As such, nbtgroups > 1 may break (but this is not yet supported without special build options).	2022-07-15 20:16:30 +02:00
Willy Tarreau	6d3c501c08	MEDIUM: fd/poller: turn update_mask to group-local IDs From now on, the FD's update_mask only refers to local thread IDs. However, there remains a limitation, in updt_fd_polling(), we temporarily have to check and set shared FDs against .thread_mask, which still contains global ones. As such, nbtgroups > 1 may break (but this is not yet supported without special build options).	2022-07-15 20:16:30 +02:00
Willy Tarreau	63022128a5	MEDIUM: fd/poller: turn polled_mask to group-local IDs This changes the signification of each bit in the polled_mask so that now each bit represents a local thread ID for the current group instead of a global thread ID. As such, all tests now apply to ltid_bit instead of tid_bit. No particular check was made to verify that the FD's tgid matches the current one because there should be no case where this is not true. A check was added in epoll's __fd_clo() to confirm it never differs unless expected (soft stop under thread isolation, or master in starting mode going to exec mode), but that doesn't prevent from doing the job: it only consists in checking in the group's threads those that are still polling this FD and to remove them. Some atomic loads were added at the various locations, and most repetitive references to polled_mask[fd].xx were turned to a local copy instead making the code much more clear.	2022-07-15 20:16:30 +02:00
Willy Tarreau	0dc1cc93b6	MAJOR: fd: grab the tgid before manipulating running We now grab a reference to the FD's tgid before manipulating the running_mask so that we're certain it corresponds to our own group (hence bits), and we drop it once we've set the bit. For now there's no measurable performance impact in doing this, which is great. The lock can be observed by perf top as taking a small share of the time spent in fd_update_events(), itself taking no more than 0.28% of CPU under 8 threads. However due to the fact that the thread groups are not yet properly spread across the pollers and the thread masks are still wrong, this will trigger some BUG_ON() in fd_insert() after a few tens of thousands of connections when threads other than those of group 1 are reached, and this is expected.	2022-07-15 20:16:30 +02:00
Willy Tarreau	c243182370	MINOR: cli/fd: show fd's tgid and refcount in "show fd" We really need to display these values now.	2022-07-15 19:58:06 +02:00
Willy Tarreau	9464bb1f05	MEDIUM: fd: add the tgid to the fd and pass it to fd_insert() The file descriptors will need to know the thread group ID in addition to the mask. This extends fd_insert() to take the tgid, and will store it into the FD. In the FD, the tgid is stored as a combination of tgid on the lower 16 bits and a refcount on the higher 16 bits. This allows to know when it's really possible to trust the tgid and the running mask. If a refcount is higher than 1 it indeed indicates another thread else might be in the process of updating these values. Since a closed FD must necessarily have a zero refcount, a test was added to fd_insert() to make sure that it is the case.	2022-07-15 19:58:06 +02:00
Willy Tarreau	512dd2dc1c	MINOR: fd: make fd_insert() apply the thread mask itself It's a bit ugly to see that half of the callers of fd_insert() have to apply all_threads_mask themselves to the bit field they're passing, because usually it comes from a listener that may have other bits set. Let's make the function apply the mask itself.	2022-07-15 19:58:06 +02:00
Willy Tarreau	8e2c0fa8e5	MINOR: fd: delete unused updates on close() After a poller's ->clo() was called to completely terminate operations on an FD, there's no reason for keeping updates on this FD, so if any updates were already programmed it would be nice if we could delete them. Tests show that __fd_clo() is called roughly half of the time with the last FD from the local update list, which possibly makes sense if a close has to appear after a polling change resulting from an incomplete read or the end of a send(). We can detect this and remove the last entry, which gives less work to do during the update() call, and eliminates most of the poll_drop_fd event reports. Note that while tempting, this must not be backported because it's only safe to be done now that fd_delete_orphan() clears the update mask as we need to be certain not to miss it: - if the update mask is kept up with no entry, we can miss future updates ; - if the update mask is cleared too fast, it may result in failure to add a shared event.	2022-07-15 19:58:06 +02:00
Willy Tarreau	35ee710ece	MEDIUM: fd/poller: make the update-list per-group The update-list needs to be per-group because its inspection is based on a mask and we need to be certain when scanning it if a mask is for the same thread or another one. Once per-group there's no doubt about it, even if the FD's polling changes, the entry remains valid. It will be needed to check the tgid though. Note that a soft-stop or pause/resume might not necessarily work here with tgroups>1, because the operation might be delivered to a thread that doesn't belong to the group and whoe update mask will not reflect one that is interesting here. We can't do better at this stage.	2022-07-15 19:57:28 +02:00
Willy Tarreau	2f36d902aa	MAJOR: fd: remove pending updates upon real close Dealing with long-lasting updates that outlive a close() is always going to be quite a problem, not because of the thread that will discover such updates late, but mostly due to the shared update_list that will have an entry on hold making it difficult to reuse it, and requiring that the fd's tgid is changed and the update_mask reset from a safe location. After careful inspection, it turns out that all our pollers that support automatic event removal upon close() do not need any extra bookkeeping, and that poll and select that use an internal representation already provide a poller->clo() callback that is already used to update the local event. As such, it is already safe to reset the update mask and to remove the event from the shared list just before the final close, because nothing remains to be done with this FD by the poller. Doing so considerably simplifies the handling of updates, which will only have to be inspected by the pollers, while the writers can continue to consider that the entries are always valid. Another benefit is that it will be possible to reduce contention on the update_list by just having one update_list per group (left to be done later if needed).	2022-07-15 19:43:10 +02:00
Willy Tarreau	15c5500b6e	MEDIUM: conn: make conn_backend_get always scan the same group We don't want to pick idle connections from another thread group, this would be very slow by forcing to share undesirable data. This patch makes sure that we start seeking from the current thread group's threads only and loops over that range exclusively. It's worth noting that the next_takeover pointer remains per-server and will bounce when multiple groups use it at the same time. But we preserve the perturbation by applying a modulo when retrieving it, so that when groups are of the same size (most common case), the index will not even change. At this time it doesn't seem worth storing one index per group in servers, but that might be an option if any contention is detected later.	2022-07-15 19:43:10 +02:00
Willy Tarreau	91a7c164b4	MINOR: task: move the niced_tasks counter to the thread group context This one is only used as a hint to improve scheduling latency, so there is no more point in keeping it global since each thread group handles its own run q	2022-07-15 19:43:10 +02:00
Willy Tarreau	b0e7712fb2	MEDIUM: task/thread: move the task shared wait queues per thread group Their migration was postponed for convenience only but now's time for having the shared wait queues per thread group and not just per process, otherwise the WQ lock uses a huge amount of CPU alone.	2022-07-15 19:43:10 +02:00
Willy Tarreau	82e378aa8a	MINOR: fd/thread: get rid of thread_mask() Since commit `d2494e048` ("BUG/MEDIUM: peers/config: properly set the thread mask") there must not remain any single case of a receiver that is bound nowhere, so there's no need anymore for thread_mask(). We're adding a test in fd_insert() to make sure this doesn't happen by accident though, but the function was removed and its rare uses were replaced with the original value of the bind_thread msak.	2022-07-15 19:43:10 +02:00
Willy Tarreau	6bdf9452c0	MINOR: cli/threads: always bind CLI to thread group 1 When using multiple groups, the stats socket starts to emit errors and it's not natural to have to touch the global section just to specify "thread 1/all". Let's pre-attach these sockets to thread group 1. This will cause errors when trying to change the group but this really is not a problem for now as thread groups are not enabled by default. This will make sure configs remain portable and may possibly be relaxed later.	2022-07-15 19:43:10 +02:00
Willy Tarreau	dcbd763fe9	MINOR: mworker/threads: limit the mworker sockets to group 1 As a side effect of commit `34aae2fd1` ("MEDIUM: mworker: set the iocb of the socketpair without using fd_insert()"), a config may now refuse to start if there are multiple groups configured because the default bind mask may span over multiple groups, and it is not possible to force it to work differently. Let's just assign thread group 1 to the master<->worker sockets so that the thread bindings automatically resolve to a single group. The same was done for the master side of the socket even if it's not used. It will avoid being forgotten in the future.	2022-07-15 19:43:10 +02:00
Willy Tarreau	5b09341c02	MEDIUM: cpu-map: replace the process number with the thread group number The principle remains the same, but instead of having a single process and ignoring extra ones, now we set the affinity masks for the respective threads of all groups. The doc was updated with a few extra examples.	2022-07-15 19:43:10 +02:00
Willy Tarreau	1b2b59bfa7	MINOR: thread: remove MAX_THREADS limitation This one is now causing difficulties during the development phase and it's going to disappear anyway, let's get rid of it.	2022-07-15 19:43:10 +02:00
Willy Tarreau	e5715bface	MEDIUM: poller: disable thread-groups for poll() and select() These old legacy pollers are not designed for this. They're still using a shared list of events for all threads, this will not scale at all, so there's no point in enabling thread-groups there. Modern systems have epoll, kqueue or event ports and do not need these ones. We arrange for failing at boot time, only when thread-groups > 1 so that existing setups will remain unaffected. If there's a compelling reason for supporting thread groups with these pollers in the future, the rework should not be too hard, it would just consume a lot of memory to have an fd_evts[] array per thread, but that is doable.	2022-07-15 19:43:10 +02:00
Willy Tarreau	b1093c6ba2	MEDIUM: poller: program the update in fd_update_events() for a migrated FD When an FD is migrated, all pollers program an update. That's useless code duplication, and when thread groups will be supported, this will require an extra round of locking just to verify the update_mask on return. Let's just program the update direction from fd_update_events() as it already does for closed FDs, this becomes more logical.	2022-07-15 19:43:10 +02:00
Willy Tarreau	1b927eb3c3	MEDIUM: proto: stop protocols under thread isolation during soft stop protocol_stop_now() is called from do_soft_stop_now() running on any thread that received the signal. The problem is that it will call some listener handlers to close the FD, resulting in an fd_delete() being called from the wrong group. That's not clean and we cannot even rely on the thread mask to show up. One interesting long-term approach could be to have kill queues for FDs, and maybe we'll need them in the long run. However that doesn't work well for listeners in this situation. Let's simply isolate ourselves during this instant. We know we'll be alone dealing with the close and that the FD will be instantly deleted since not in use by any other thread. It's not the cleanest solution but it should last long enough without causing trouble.	2022-07-15 19:43:10 +02:00
Willy Tarreau	7aa41196cf	MEDIUM: debug/threads: make the lock debugging take tgroups into account Since we have to use masks to verify owners/waiters, we have no other option but to have them per group. This definitely inflates the size of the locks, but this is only used for extreme debugging anyway so that's not dramatic. Thus as of now, all masks in the lock stats are local bit masks, derived from ti->ltid_bit. Since at boot ltid_bit might not be set, we just take care of this situation (since some structs are initialized under look during boot), and use bit 0 from group 0 only.	2022-07-15 19:41:26 +02:00
Willy Tarreau	4d9888ca69	CLEANUP: fd: get rid of the __GET_{NEXT,PREV} macros They were initially made to deal with both the cache and the update list but there's no cache anymore and keeping them for the update list adds a lot of obfuscation that is really not desired. Let's get rid of them now. Their purpose was simply to get a pointer to fdtab[fd].update.{,next,prev} in order to perform atomic tests and modifications. The offset passed in argument to the functions (fd_add_to_fd_list() and fd_rm_from_fd_list()) was the offset of the ->update field in fdtab, and as it's not used anymore it was removed. This also removes a number of casts, though those used by the atomic ops have to remain since only scalars are supported.	2022-07-15 19:41:26 +02:00
Willy Tarreau	740038c8b9	MINOR: listener/config: make "thread" always support up to LONGBITS The difference is subtle but in one place there was MAXTHREADS and this will not work anymore once it goes over 64.	2022-07-15 19:41:26 +02:00
Willy Tarreau	acd644197f	MEDIUM: config: remove the "process" keyword on "bind" lines It was deprecated, marked for removal in 2.7 and was already emitting a warning, let's get rid of it. Note that we've kept the keyword detection to suggest to use "thread" instead.	2022-07-15 19:41:26 +02:00
Willy Tarreau	94f763b5e4	MEDIUM: config: remove deprecated "bind-process" directives from frontends This was already causing a deprecation warning and was marked for removal in 2.7, now it happens. An error message indicates this doesn't exist anymore.	2022-07-15 19:41:26 +02:00
Willy Tarreau	91f7a1af34	CLEANUP: applet: remove the obsolete command context from the appctx The "ctx" and "st2" parts in the appctx were marked for removal in 2.7 and were emulated using memcpy/memset etc for possible external code. Let's remove this now.	2022-07-15 19:41:26 +02:00
Willy Tarreau	9a7fa90239	MINOR: cli/activity: add a thread number argument to "show activity" The output of "show activity" can be so large that the output is visually unreadable on a screen. Let's add an option to filter on the desired column (actually the thread number), use "0" to report only the first column (aggregated/sum/avg), and use "-1", the default, for the normal detailed dump.	2022-07-15 19:41:26 +02:00
Willy Tarreau	dadf00e226	DEBUG: cli: add a new "debug dev deadlock" expert command This command will create the requested number of tasks competing on a lock, resulting in triggering the watchdog and crashing the process. This will help stress the watchdog and inspect the lock debugging parts.	2022-07-15 19:41:26 +02:00
Willy Tarreau	dd75b64cdf	MINOR: cli/streams: show a stream's tgid next to its thread ID We now display both the global thread ID and the tgid/ltid pair so that it's easier to match it with the FD.	2022-07-15 19:41:26 +02:00
Willy Tarreau	f0c86ddfe8	BUG/MEDIUM: debug: fix parallel thread dumps again The previous attempt to fix thread dumps in commit `672972604` ("BUG/MEDIUM: debug: fix possible hang when multiple threads dump at once") still had some shortcomings. Sometimes parallel dumps are jerky essentially due to the way that threads synchronize on startup and end. In addition the risk of waiting forever for a stopped thread exists, and panics happening in parallel to thread dumps are not more reliable either. This commit revisits the state transitions so that all threads may request a dump in parallel, that all of them wait for each other in the handler, and that one thread is responsible for counting every other and checking that the total matches the number of active threads. Then for stopping there's a finishing phase that all threads wait for so that none quits this area too early. Given that we now know the number of participants to the dump, we can let them each decrement the counter when leaving so that another dump may only start after the last participant has completely left. Now many thread dumps in parallel are running fine, so do panics. No backport is needed as this was the result of the changes for thread groups.	2022-07-15 19:41:26 +02:00
Willy Tarreau	55433f9b34	BUG/MINOR: debug: enter ha_panic() only once Some panic dumps are mangled or truncated due to the watchdog firing at the same time on multiple threads and calling ha_panic() simultaneously. What may happen in this case is that the second one waits for the first one to finish but as soon as it's done the second one resets the buffer and dumps again, sometimes resetting the first one's dump. Also the first one's abort() may trigger while the second one is currently dumping, resulting in a full dump followed by a truncated one, leading to confusion. Sometimes some lines appear in the middle of a dump as well. It doesn't happen often and is easier to trigger by causing massive deadlocks. There's no reason for the process to resist to a panic, so we can safely add a counter and no nothing on subsequent calls. Ideally we'd wait there forever but as this may happen inside a signal handler (e.g. watchdog), it doesn't always work, so the easiest thing to do is to return so that the thread is interrupted as soon as possible and brought to the debug handler to be dumped. This should be backported, at least to 2.6 and possibly to older versions as well.	2022-07-15 19:41:26 +02:00
Willy Tarreau	f15c75a2d3	BUG/MINOR: thread: use the correct thread's group in ha_tkillall() In ha_tkillall(), the current thread's group was used to check for the thread being running instead of using the target thread's group mask. Most of the time it would not have any effect unless some groups are uneven where it can lead to incomplete thread dumps for example. No backport is needed, this is purely 2.7.	2022-07-15 19:41:26 +02:00
Willy Tarreau	52f238d326	BUG/MEDIUM: cli/threads: make "show threads" more robust on applets Running several concurrent "show threads" in loops might occasionally cause a segfault when trying to retrieve the stream from appctx_sc() which may be null while the applet is finishing. It's not easy to reproduce, it requires 3-5 sessions in parallel for about a minute or so. The appctx_sc must be checked before passing it to sc_strm(). This must be backported to 2.6 which also has the bug.	2022-07-15 19:41:26 +02:00
Willy Tarreau	9b0f0d146f	BUG/MINOR: threads: produce correct global mask for tgroup > 1 In thread_resolve_group_mask(), if a global thread number is passed and it belongs to a group greater than 1, an incorrect shift resulted in shifting that ID again which made it appear nowhere or in a wrong group possibly. The bug was introduced in 2.5 with commit `627def9e5` ("MINOR: threads: add a new function to resolve config groups and masks") though the groups only starts to be usable in 2.7, so there is no impact for this bug, hence no backport is needed.	2022-07-15 19:41:26 +02:00
Amaury Denoyelle	114c9c87ce	MINOR: h3: implement graceful shutdown with GOAWAY Implement graceful shutdown as specified in RFC 9114. A GOAWAY frame is generated with stream ID to indicate range of processed requests. This process is done via the release app protocol operation. The MUX is responsible to emit the generated GOAWAY frame after app release. A CONNECTION_CLOSE will be emitted once there is no unacknowledged STREAM frames.	2022-07-15 15:56:13 +02:00
Amaury Denoyelle	d701039773	MINOR: h3: store control stream in h3c Store a reference to the HTTP/3 control stream in h3c context. This will be useful to implement GOAWAY emission without having to store the control stream ID on opening.	2022-07-15 15:56:13 +02:00
Amaury Denoyelle	a154dc0290	MINOR: mux-quic: send one last time before release Call qc_send() on qc_release(). This is mostly useful for application protocol with a connection closing procedure. Most notably, this will be useful to implement HTTP/3 GOAWAY emission.	2022-07-15 15:56:13 +02:00
Amaury Denoyelle	c49d5d1a4b	CLEANUP: mux-quic: move qc_release() This change is purely cosmetic. qc_release() function is moved just before qc_io_cb(). It's cleaner as it brings it closer where it is used. More importantly, this will be required to be able to use it in qc_send() function.	2022-07-15 15:56:13 +02:00
Amaury Denoyelle	240b1b108b	MEDIUM: quic: send CONNECTION_CLOSE on released MUX Send a CONNECTION_CLOSE if the MUX has been released and all STREAM data are acknowledged. This is useful to prevent a client from trying to use a connection which have the upper layer closed. To implement this a new function qc_check_close_on_released_mux() has been added. It is called on QUIC MUX release notification and each time a qc_stream_desc has been released. This commit is associated with the previous one : MINOR: mux-quic/h3: schedule CONNECTION_CLOSE on app release Both patches are required to prevent the risk of browsers stuck on webpage loading if MUX has been released. On CONNECTION_CLOSE reception, the client will reopen a new QUIC connection.	2022-07-15 15:56:13 +02:00
Amaury Denoyelle	069288b4c0	MINOR: mux-quic/h3: prepare CONNECTION_CLOSE on release When MUX is released, a CONNECTION_CLOSE frame should be emitted. This will ensure that the client does not use anymore a half-dead connection. App protocol layer is responsible to provide the error code via release callback. For HTTP/3 NO_ERROR is used as specified in RFC 9114. If no release callback is provided, generic QUIC NO_ERROR code is used. Note that a graceful shutdown is used : quic_conn must emit CONNECTION_CLOSE frame when possible. This will be provided in another patch. This change should limit the risk of browsers stuck on webpage loading if MUX has been released. On CONNECTION_CLOSE reception, the client will reopen a new QUIC connection.	2022-07-15 15:20:33 +02:00
Amaury Denoyelle	d666d740d2	MINOR: mux-quic: support app graceful shutdown Adjust qcc_emit_cc_app() to allow the delay of emission of a CONNECTION_CLOSE. This will only set the error code but the quic-conn layer is not flagged for immediate close. The quic-conn will be responsible to shut the connection when deemed suitable. This change will allow to implement application graceful shutdown, such as HTTP/3 with GOAWAY emission. This will allow to emit closing frames on MUX release. Once all work is done at the lower layer, the quic-conn should emit a CONNECTION_CLOSE with the registered error code.	2022-07-15 15:06:59 +02:00
Amaury Denoyelle	57e6db7021	MINOR: quic: define a generic QUIC error type Define a new structure quic_err to abstract a QUIC error type. This allows to easily differentiate a transport and an application error code. This simplifies error transmission from QUIC MUX and H3 layers. This new type is defined in quic_frame module. It is used to replace <err_code> field in <quic_conn>. QUIC_FL_CONN_APP_ALERT flag is removed as it is now useless. Utility functions are defined to be able to quickly instantiate transport, tls and application errors.	2022-07-15 14:57:49 +02:00
Amaury Denoyelle	72d86509f1	BUG/MINOR: quic: fix closing state on NO_ERROR code sent Reception is disabled as soon as a CONNECTION_CLOSE emission is required. An early return is done on qc_lstnr_pkt_rcv() to implement this. This condition is not functional if the error code sent is NO_ERROR (0x00). To fix this, check the quic-conn flags instead of the error code. Currently this bug has no impact has NO_ERROR emission is not used. This can be backported up to 2.6.	2022-07-13 15:33:15 +02:00
Willy Tarreau	672972604f	BUG/MEDIUM: debug: fix possible hang when multiple threads dump at once A bug in the thread dumper was introduced by commit `00c27b50c` ("MEDIUM: debug: make the thread dumper not rely on a thread mask anymore"). If two or more threads try to trigger a thread dump exactly at the same time, the second one may loop indefinitely trying to set the value to 1 while the other ones will wait for it to finish dumping before leaving. This is a consequence of a logic change using thread numbers instead of a thread mask, as threads do not need to see all other ones there anymore. No backport is needed, this is only for 2.7.	2022-07-13 09:03:02 +02:00
Amaury Denoyelle	a5b5075211	MEDIUM: mux-quic: implement STOP_SENDING handling Implement support for STOP_SENDING frame parsing. The stream is resetted as specified by RFC 9000. This will automatically interrupt all future send operation in qc_send(). A RESET_STREAM will be sent with the code extracted from the original STOP_SENDING frame.	2022-07-11 16:45:04 +02:00
Amaury Denoyelle	843a1196b3	MEDIUM: mux-quic: implement RESET_STREAM emission Implement functions to be able to reset a stream via RESET_STREAM. If needed, a qcs instance is flagged with QC_SF_TO_RESET to schedule a stream reset. This will interrupt all future send operations. On stream emission, if a stream is flagged with QC_SF_TO_RESET, a RESET_STREAM frame is generated and emitted to the transport layer. If this operation succeeds, the stream is locally closed. If upper layer is instantiated, error flag is set on it.	2022-07-11 16:45:04 +02:00
Amaury Denoyelle	20d1f84ce4	MINOR: mux-quic: use stream states to mark as detached Adjust condition to detach a qcs instance : if the stream is not locally close it is not directly free. This should improve stream closing by ensuring that either FIN or a RESET_STREAM is sent before destroying it.	2022-07-11 16:41:10 +02:00
Amaury Denoyelle	38e6006da1	MINOR: mux-quic: define basic stream states Implement a basic state machine to represent stream lifecycle. By default a stream is idle. It is marked as open when sending or receiving the first data on a stream. Bidirectional streams has two states to represent the closing on both receive and send channels. This distinction does not exists for unidirectional streams which passed automatically from open to close state. This patch is mostly internal and has a limited visible impact. Some behaviors are slightly updated : * closed streams are garbage collected at the start of io handler * send operation is interrupted if a stream is close locally Outside of this, there is no functional change. However, some additional BUG_ON guards are implemented to ensure that we do not conduct invalid operation on a stream. This should strengthen the code safety. Also, stream states are displayed on trace which should help debugging.	2022-07-11 16:37:21 +02:00
Amaury Denoyelle	b68559a9aa	MINOR: mux-quic: support stream opening via MAX_STREAM_DATA MAX_STREAM_DATA can be used as the first frame of a stream. In this case, the stream should be opened, if it respects flow-control limit. To implement this, simply replace plain lookup in stream tree by qcc_get_qcs() at the start of the parsing function. This automatically takes care of opening the stream if not already done. As specified by RFC 9000, if MAX_STREAM_DATA is receive for a receive-only stream, a STREAM_STATE_ERROR connection error is emitted.	2022-07-11 16:24:03 +02:00
Amaury Denoyelle	57161b7d0c	MINOR: mux-quic: do not ack STREAM frames on unrecoverable error Improve return path for qcc_recv() on STREAM parsing. It returns 0 on success. On error, a non-zero value is returned which indicates to the caller that the packet containing the frame should not be acknowledged. When qcc_recv() generates a CONNECTION_CLOSE or RESET_STREAM, either directly or via qcc_get_qcs(), an error is returned which ensure that no acknowledgement is generated. This required an adjustment on qcc_get_qcs() API which now returns a success/error code. The stream instance is returned via a new out argument.	2022-07-11 16:24:03 +02:00
Amaury Denoyelle	5fbb8691d4	MINOR: mux-quic: filter send/receive-only streams on frame parsing Extend the function qcc_get_qcs() to be able to filter send/receive-only unidirectional streams. A connection error STREAM_STATE_ERROR is emitted if this new filter does not match. This will be useful when various frames handlers are converted with qcc_get_qcs(). Depending on the frame type, it will be easy to filter on the forbidden stream types as specified in RFC 9000.	2022-07-11 16:24:03 +02:00
Amaury Denoyelle	4561f84ad4	MINOR: mux-quic: implement qcs_alert() Implement a simple function to notify a possible subscriber or wake up the upper layer if a special condition happens on a stream. For the moment, this is only used to replace identical code in qc_wake_some_streams().	2022-07-11 16:24:03 +02:00
Amaury Denoyelle	392e94e985	MINOR: mux-quic: add traces on frame parsing functions Add traces for parsing functions for MAX_DATA and MAX_STREAM_DATA.	2022-07-11 16:24:03 +02:00
Amaury Denoyelle	c1a6dfd477	MINOR: mux-quic: rename stream purge function Rename qc_release_detached_streams() to qc_purge_streams(). The aim is to have a more generic name. It's expected to complete this function to add other criteria to purge dead streams. Also the function documentation has been corrected. It does not return a number of streams. Instead it is a boolean value, to true if at least one stream was released.	2022-07-11 16:24:03 +02:00
Amaury Denoyelle	b143723411	REORG: mux-quic: rename stream initialization function Rename both qcc_open_stream_local/remote() functions to qcc_init_stream_local/remote(). This change is purely cosmetic. It will reduces the ambiguity with the soon to be implemented OPEN states for QCS instances.	2022-07-11 16:24:03 +02:00
Amaury Denoyelle	e53b489826	BUG/MEDIUM: mux-quic: fix server chunked encoding response QUIC MUX was not able to correctly deal with server response using chunked transfer-encoding. All data will be transfered correctly to the client but the FIN bit is missing. The transfer will never stop as the client will wait indefinitely for the FIN bit. This bug happened because the HTX message representing a chunked encoded payload contains a final empty block with the EOM flag. However, emission is skipped by QUIC MUX if there is no data to transfer. To fix this, the condition was completed to ensure that there is no need to send the FIN signal. If this is false, data emission will proceed even if there is no data : this will generate an empty QUIC STREAM frame with FIN set which will mark the end of the transfer. To ensure that a FIN STREAM frame is sent only one time, QC_SF_FIN_STREAM is resetted on send confirmation from the transport in qcc_streams_sent_done(). This bug was reproduced when dealing with chunked transfer-encoding response for the HTTP server. This must be backported up to 2.6.	2022-07-11 16:21:52 +02:00
Willy Tarreau	a88e8bf428	BUILD: http: silence an uninitialized warning affecting gcc-5 When building with gcc-5, one can see this warning: src/http_fetch.c: In function 'smp_fetch_meth': src/http_fetch.c:356:6: warning: 'htx' may be used uninitialized in this function [-Wmaybe-uninitialized] sl = http_get_stline(htx); ^ It's wrong since the only way to reach this code is to have met the same condition a few lines before and initialized the htx variable. The reason in fact is that the same test happens on different variables of distinct types, so the compiler possibly doesn't know that the condition is the same. Newer gcc versions do not have this problem. Let's just move the assignment earlier and have the exact same test, as it's sufficient to shut this up. This may have to be backported to 2.6 since the code is the same there.	2022-07-10 14:13:48 +02:00
Willy Tarreau	0d023774bf	MEDIUM: epoll: don't synchronously delete migrated FDs Between 1.8 and 1.9 commit `d9e7e36c6` ("BUG/MEDIUM: epoll/threads: use one epoll_fd per thread") split the epoll poller to use one poller per thread (and this was backported to 1.8). This patch added a call to epoll_ctl(DEL) on return from the I/O handler as a safe way to deal with a detected thread migration when that code was still quite fragile. One aspect of this choice was that by then we wanted to maintain support for the rare old bogus epoll implementations that failed to remove events on close(), so risking to lose the event was not an option. Later in 2.5, commit `200bd50b7` ("MEDIUM: fd: rely more on fd_update_events() to detect changes") changed the code to perform most of the operations inside fd_update_events(), but it maintained that oddity, to the point that strictly all pollers except epoll now just add an update to be dealt with at the next round. This approach is much more efficient, because under load and server-side connection reuse, it's perfectly possible for a thread to see the same FD several times in a poll loop, the first time to relinquish it after a migration, then the other thread makes a request, gets its response, and still during the same loop for the first one, grabbing an idle connection to send a request and wait for a response will program a new update on this FD. By using a synchronous epoll_ctl(DEL), we effectively lose the opportunity to aggregate certain changes in the same update. Some tests performed locally with 8 threads and one server show that on average, by using an update instead of a synchronous call, we reduce the number of epoll_ctl() calls by 25-30% (under low loads it will probably not change anything). So this patch implements the same method for all pollers and replaces the synchronous epoll_ctl() with an update.	2022-07-10 14:13:48 +02:00
Christopher Faulet	372b38f935	BUG/MEDIUM: mux-h1: Handle connection error after a synchronous send Since commit `d1480cc8` ("BUG/MEDIUM: stream-int: do not rely on the connection error once established"), connection errors are not handled anymore by the stream-connector once established. But it is a problem for the H1 mux when an error occurred during a synchronous send in h1_snd_buf(). Because in this case, the connction error is just missed. It leads to a session leak until a timeout is reached (client or server). To fix the bug, the connection flags are now checked in h1_snd_buf(). If there is an error, it is reported to the stconn level by setting SF_FL_ERROR flags. But only if there is no pending data in the input buffer. This patch should solve the issue #1774. It must be backported as far as 2.2.	2022-07-08 16:37:31 +02:00
Christopher Faulet	52fc0cbaad	BUG/MEDIUM: http-ana: Don't wait to have an empty buf to switch in TUNNEL state When we want to establish a tunnel on a side, we wait to have flush all data from the buffer. At this stage the other side is at least in DONE state. But there is no real reason to wait. We are already in DONE state on its side. So all the HTTP message was already forwarded or planned to be forwarded. Depending on the scheduling if the mux already started to transfer tunneled data, these data may block the switch in TUNNEL state and thus block these data infinitly. This bug exists since the early days of HTX. May it was mandatory but today it seems useless. But I honestly don't remember why this prerequisite was added. So be careful during the backports. This patch should be backported with caution. At least as far as 2.4. For 2.2 and 2.0, it seems to be mandatory too. But a review must be performed.	2022-07-08 16:37:31 +02:00
Christopher Faulet	5966e40641	BUG/MINOR: mux-h1: Be sure to commit htx changes in the demux buffer When a buffer area is casted to an htx message, depending on the method used, the underlying buffer may be updated or not. The htxbuf() function does not change the buffer state. We assume the buffer was already prepared to store an htx message. htx_from_buf() on its side, updates the buffer. With the first function, we only need to commit changes to the underlying buffer if the htx message is changed. With last one, we must always commit the changes. The idea is to be sure to keep non-empty HTX messages while an empty message must be lead to an empty buffer after commit. All that said because in h1_process_demux(), the changes is not always committed as expected. When the demux is blocked, we just leave the function. So it is possible to have an empty htx message stored in a buffer that appears non-empty. It is not critical, but the buffer cannot be released in this state. And we should always release an empty buffer. This patch must be backported as far as 2.4.	2022-07-08 16:37:31 +02:00
William Lallemand	a46a99e98c	MEDIUM: mworker/systemd: send STATUS over sd_notify The sd_notify API is not able to change the "Active:" line in "systemcl status". However a message can still be displayed on a "Status: " line, even if the service is still green and "active (running)". When startup succeed the Status will be set to "Ready.", upon a reload it will be set to "Reloading Configuration." If the configuration succeed "Ready." again. However if the reload failed, it will be set to "Reload failed!". Keep in mind that the "Active:" line won't change upon a reload failure, and will still be green.	2022-07-07 14:48:46 +02:00
Christopher Faulet	12f6dbb863	BUG/MEDIUM: http-fetch: Don't fetch the method if there is no stream The "method" sample fetch does not perform any check on the stream existence before using it. However, for errors triggered at the mux level, there is no stream. When the log message is formatted, this sample fetch must fail. It must also fail when it is called from a health-check. This patch must be backported as far as 2.4.	2022-07-07 09:35:58 +02:00
Christopher Faulet	d1d983fb12	MINOR: http-htx: Use new HTTP functions for the scheme based normalization Use http_get_host_port() and http_is_default_port() functions to perform the scheme based normalization.	2022-07-07 09:35:58 +02:00
Christopher Faulet	3f5fbe9407	BUG/MEDIUM: h1: Improve authority validation for CONNCET request From time to time, users complain to get 400-Bad-request responses for totally valid CONNECT requests. After analysis, it is due to the H1 parser performs an exact match between the authority and the host header value. For non-CONNECT requests, it is valid. But for CONNECT requests the authority must contain a port while it is often omitted from the host header value (for default ports). So, to be sure to not reject valid CONNECT requests, a basic authority validation is now performed during the message parsing. In addition, the host header value is normalized. It means the default port is removed if possible. This patch should solve the issue #1761. It must be backported to 2.6 and probably as far as 2.4.	2022-07-07 09:35:58 +02:00
Christopher Faulet	ca7218aaf0	MINOR: http: Add function to detect default port http_is_default_port() can be used to test if a port is a default HTTP/HTTPS port. A scheme may be specified. In this case, it is used to detect defaults ports, 80 for "http://" and 443 for "https://". Otherwise, with no scheme, both are considered as default ports.	2022-07-06 17:54:03 +02:00
Christopher Faulet	658f971621	MINOR: http: Add function to get port part of a host http_get_host_port() function can be used to get the port part of a host. It will be used to get the port of an uri authority or a host header value. This function only look for a port starting from the end of the host. It is the caller responsibility to call it with a valid host value. An indirect string is returned.	2022-07-06 17:54:03 +02:00
Christopher Faulet	0eab050b04	BUG/MINOR: http-htx: Fix scheme based normalization for URIs wih userinfo The scheme based normalization is not properly handled the URI's userinfo, if any. First, the authority parser is not called with "no_userinfo" parameter set. Then it is skipped from the URI normalization. This patch must be backported as far as 2.4.	2022-07-06 17:54:02 +02:00
William Lallemand	1d93217a05	BUG/MINOR: peers: fix possible NULL dereferences at config parsing Patch `49f6f4b` ("BUG/MEDIUM: peers: fix segfault using multiple bind on peers sections") introduced possible NULL dereferences when parsing the peers configuration. Fix the issue by checking the return value of bind_conf_uniq_alloc(). This patch should be backported as far as 2.0.	2022-07-06 14:40:11 +02:00
Willy Tarreau	ad92fdf196	CLEANUP: thread: also remove a thread's bit from stopping_threads on stop As much as possible we should take care of not leaving bits from stopped threads in shared thread masks. It can avoid issues like the previous fix and will also make debugging less confusing.	2022-07-06 10:19:46 +02:00
Willy Tarreau	f34a3fa33d	BUG/MEDIUM: thread: mask stopping_threads with threads_enabled when checking it When soft-stopping, there's a comparison between stopping_threads and threads_enabled to make sure all threads are stopped, but this is not correct and is racy since the threads_enabled bit is removed when a thread is stopped but not its stopping_threads bit. The consequence is that depending on timing, when stopping, if the first stopping thread is fast enough to remove its bit from threads_enabled, the other threads will see that stopping_threads doesn't match threads_enabled anymore and will wait forever. As such the mask must be applied to stopping_threads during the test. This issue was introduced in recent commit `ef422ced9` ("MEDIUM: thread: make stopping_threads per-group and add stopping_tgroups"), no backport is needed.	2022-07-06 10:19:46 +02:00
Christopher Faulet	4c3d3d2a68	BUG/MINOR: http-act: Properly generate 103 responses when several rules are used When several "early-hint" rules are used, we try, as far as possible, to merge links into the same 103-early-hints response. However, it only works if there is no ACLs. If a "early-hint" rule is not executed an invalid response is generated. the EOH block or the start-line may be missing, depending on the rule order. To fix the bug, we use the transaction status code. It is unused at this stage. Thus, it is set to 103 when a 103-early-hints response is in progress. And it is reset when the response is forwarded. In addition, the response is forwarded if the next rule is an "early-hint" rule with an ACL. This way, the response is always valid. This patch must be backported as far as 2.2.	2022-07-06 09:37:43 +02:00
Christopher Faulet	4c8e58def6	BUG/MINOR: http-check: Preserve headers if not redefined by an implicit rule When an explicit "http-check send" rule is used, if it is the first one, it is merge with the implicit rule created by "option httpchk" statement. The opposite is also true. Idea is to have only one send rule with the merged info. It means info defined in the second rule override those defined in the first one. However, if an element is not defined in the second rule, it must be ignored, keeping this way info from the first rule. It works as expected for the method, the uri and the request version. But it is not true for the header list. For instance, with the following statements, a x-forwarded-proto header is added to healthcheck requests: option httpchk http-check send meth GET hdr x-forwarded-proto https while by inverting the statements, no extra headers are added: http-check send meth GET hdr x-forwarded-proto https option httpchk Now the old header list is overriden if the new one is not empty. This patch should fix the issue #1772. It must be backported as far as 2.2.	2022-07-06 09:35:13 +02:00
Christopher Faulet	f0196f4f71	CLEANUP: bwlim: Set pointers to NULL when memory is released Calls to free() are replaced by ha_free(). And otherwise, the pointers are explicitly set to NULL after a release. There is no issue here but it could help debugging sessions.	2022-07-06 09:34:54 +02:00
Willy Tarreau	d2494e0489	BUG/MEDIUM: peers/config: properly set the thread mask The peers didn't have their bind_conf thread mask nor group set, because they're still not part of the global proxy list. Till 2.6 it seems it does not have any visible impact, since most listener-oriented operations pass through thread_mask() which detects null masks and turns them to all_threads_mask. But starting with 2.7 it becomes a problem as won't permit these null masks anymore. This patch duplicates (yes, sorry) the loop that applies to the frontend's bind_conf, though it is simplified (no sharding, etc). As the code is right now, it simply seems impossible to trigger the second (and largest) part of the check when leaving thread_resolve_group_mask() on success, so it looks like it might be removed. No backport is needed, unless a report in 2.6 or earlier mentions an issue with a null thread_mask.	2022-07-05 19:10:26 +02:00
Willy Tarreau	8d158132bd	BUG/MINOR: peers/config: always fill the bind_conf's argument Some generic frontend errors mention the bind_conf by its name as "bind '%s'", but if this is used on peers "bind" lines it shows "(null)" because the argument is set to NULL in the call to bind_conf_uniq_alloc() instead of passing the argument. Fortunately that's trivial to fix. This may be backported to older versions.	2022-07-05 19:06:47 +02:00
Amaury Denoyelle	bf91e3922b	MINOR: mux-quic: emit FINAL_SIZE_ERROR on invalid STREAM size Add a check on stream size when the stream is in state Size Known. In this case, a STREAM frame cannot change the stream size. If this is not respected, a CONNECTION_CLOSE with FINAL_SIZE_ERROR will be emitted as specified in the RFC 9000.	2022-07-05 16:44:01 +02:00
Amaury Denoyelle	3f39b40fe0	MINOR: mux-quic: rename qcs flag FIN_RECV to SIZE_KNOWN Rename QC_SF_FIN_RECV to the more generic name QC_SF_SIZE_KNOWN. This better align with the QUIC RFC 9000 which uses the "Size Known" state definition. This change is purely cosmetic.	2022-07-05 16:18:27 +02:00
Amaury Denoyelle	a509ffb505	MEDIUM: mux-quic: refactor streams opening Review the whole API used to access/instantiate qcs. A public function qcc_open_stream_local() is available to the application protocol layer. It allows to easily opening a local stream. The ID is automatically attributed to the next one available. For remote streams, qcc_open_stream_remote() has been implemented. It will automatically take care of allocating streams in a linear way according to the ID. This function is called via qcc_get_qcs() which can be used for each qcc_recv*() operations. For the moment, it is only used for STREAM frames via qcc_recv(), but soon it will be implemented for other frames types which can also be used to open a new stream. qcs_new() and qcs_free() has been restricted to the MUX QUIC only as they are now reserved for internal usage. This change is a pure refactoring and should not have any noticeable impact. It clarifies the developer intent and help to ensure that a stream is not automatically opened when not desired.	2022-07-05 16:18:27 +02:00
Amaury Denoyelle	3abeb57909	MINOR: mux-quic: implement accessor for sedesc Implement a function <qcs_sc> to easily access to the stconn associated with a QCS. This takes care of qcs.sd which may be NULL, for example for unidirectional streams. It is expected that in the future when implementing STOP_SENDING/RESET_STREAM, stconn must be notify about the event. This accessor will allow to easily test if the stconn is instantiated or not.	2022-07-05 11:22:15 +02:00
Amaury Denoyelle	a441ec9c7a	CLEANUP: mux-quic: do not export qc_get_ncbuf qc_get_ncbuf() is only used internally : thus its prototype in QUIC MUX include is not required.	2022-07-05 11:06:52 +02:00
William Lallemand	2ee490f613	CLEANUP: mworker: rename mworker_pipe to mworker_sockpair The function mworker_pipe_register_per_thread() is called this way because the master first used pipes instead of socketpairs. Rename mworker_pipe_register_per_thread() to mworker_sockpair_register_per_thread() in order to be more consistent. Also update a comment inside the function.	2022-07-05 09:06:04 +02:00
William Lallemand	34aae2fd12	MEDIUM: mworker: set the iocb of the socketpair without using fd_insert() The worker was previously changing the iocb of the socketpair in the worker by mworker_accept_wrapper(). However, it was done using fd_insert() instead of changing directly the callback in the fdtab[].iocb pointer. This patch cleans up this by part by removing fd_insert(). It also stops setting tid_bit on the thread mask, the socketpair will be handled by any thread from now.	2022-07-05 05:18:51 +02:00
Willy Tarreau	7509ec369a	MINOR: proxy: use tg->threads_enabled in hard_stop() to detect stopped threads Let's rely on tg->threads_enabled there to detect running threads. We should probably have a dedicated function for this in order to simplify the code and avoid the risk of using the wrong group ID.	2022-07-04 14:09:39 +02:00
Willy Tarreau	24cfc9f76e	BUG/MEDIUM: thread: check stopping thread against local bit and not global one Commit `ef422ced9` ("MEDIUM: thread: make stopping_threads per-group and add stopping_tgroups") moved the stopping_threads mask to per-group, but one test in the loop preserved its global value instead, resulting in stopping threads never sleeping on stop and eating 100% CPU until all were stopped. No backport is needed.	2022-07-04 14:09:39 +02:00
Willy Tarreau	291f6ff885	BUG/MEDIUM: threads: fix incorrect thread group being used on soft-stop Commit `377e37a80` ("MINOR: tinfo: add the mask of enabled threads in each group") forgot -1 on the tgid, thus the groups was not always correctly tested, which is visible only when running with more than one group. No backport is needed.	2022-07-04 13:37:31 +02:00
Willy Tarreau	89ed89e895	BUILD: debug: re-export thread_dump_state Building with threads and without thread dump (e.g. macos, freebsd) warns that thread_dump_state is unused. This happened in fact with recentcommit `1229ef312` ("MINOR: wdt: do not rely on threads_to_dump anymore"). The solution would be to mark it unused, but after a second thought, it can be convenient to keep it exported to help debug crashes, so let's export it again. It's just not referenced in include files since it's not needed outside.	2022-07-01 21:18:03 +02:00
Willy Tarreau	039972b4e5	BUILD: debug: fix build issue on clang with previous commit Since the thread_dump_state type changed to uint, the old value in the CAS needs to be the same as well.	2022-07-01 19:37:42 +02:00
Willy Tarreau	00c27b50c0	MEDIUM: debug: make the thread dumper not rely on a thread mask anymore The thread mask is too short to dump more than 64 bits. Thus here we're using a different approach with two counters, one for the next thread ID to dump (which always exists, as it's looked up), and the second one for the number of threads done dumping. This allows to dump threads in ascending order then to let them wait for all others to be done, then to leave without the risk of an overlapping dump until the done count is null again. This allows to remove threads_to_dump which was the last non-FD variable using a global thread mask.	2022-07-01 19:31:39 +02:00
Willy Tarreau	1229ef312d	MINOR: wdt: do not rely on threads_to_dump anymore This flag is not needed anymore as we're already marking the waiting threads as harmless, thus the thread's bit is already covered by this information. The variable was unexported.	2022-07-01 19:26:35 +02:00
Willy Tarreau	f7afdd910b	MINOR: debug: mark oneself harmless while waiting for threads to finish The debug_handler() function waits for other threads to join, but does not mark itself as harmless, so if at the same time another thread tries to isolate, this may deadlock. In practice this does not happen as the signal is received during epoll_wait() hence under harmless mode, but it can possibly arrive under other conditions. In order to improve this, while waiting for other threads to join, we're now marking the current thread as harmless, as it's doing nothing but waiting for the other ones. This way another harmless waiter will be able to proceed. It's valid to do this since we're not doing anything else in this loop. One improvement could be to also check for the thread being idle and marking it idle in addition to harmless, so that it can even release a full isolation requester. But that really doesn't look worth it.	2022-07-01 19:26:35 +02:00
Willy Tarreau	a2b8ed4b44	MINOR: thread: add is_thread_harmless() to know if a thread already is harmless The harmless status is not re-entrant, so sometimes for signal handling it can be useful to know if we're already harmless or not. Let's add a function doing that, and make the debugger use it instead of manipulating the harmless mask.	2022-07-01 19:26:35 +02:00
Willy Tarreau	598cf3f22e	MAJOR: threads: change thread_isolate to support inter-group synchronization thread_isolate() and thread_isolate_full() were relying on a set of thread masks for all threads in different states (rdv, harmless, idle). This cannot work anymore when the number of threads increases beyond LONGBITS so we need to change the mechanism. What is done here is to have a counter of requesters and the number of the current isolated thread. Threads which want to isolate themselves increment the request counter and wait for all threads to be marked harmless (or idle) by scanning all groups and watching the respective masks. This is possible because threads cannot escape once they discover this counter, unless they also want to isolate and possibly pass first. Once all threads are harmless, the requesting thread tries to self-assign the isolated thread number, and if it fails it loops back to checking all threads. If it wins it's guaranted to be alone, and can drop its harmless bit, so that other competing threads go back to the loop waiting for all threads to be harmless. The benefit of proceeding this way is that there's very little write contention on the thread number (none during work), hence no cache line moves between caches, thus frozen threads do not slow down the isolated one. Once it's done, the isolated thread resets the thread number (hence lets another thread take the place) and decrements the requester count, thus possibly releasing all harmless threads. With this change there's no more need for any global mask to synchronize any thread, and we only need to loop over a number of groups to check 64 threads at a time per iteration. As such, tinfo's threads_want_rdv could be dropped. This was tested with 64 threads spread into 2 groups, running 64 tasks (from the debug dev command), 20 "show sess" (thread_isolate()), 20 "add server blah/blah" (thread_isolate()), and 20 "del server blah/blah" (thread_isolate_full()). The load remained very low (limited by external socat forks) and no stuck nor starved thread was found.	2022-07-01 19:15:15 +02:00
Willy Tarreau	ef422ced91	MEDIUM: thread: make stopping_threads per-group and add stopping_tgroups Stopping threads need a mask to figure who's still there without scanning everything in the poll loop. This means this will have to be per-group. And we also need to have a global stopping groups mask to know what groups were already signaled. This is used both to figure what thread is the first one to catch the event, and which one is the first one to detect the end of the last job. The logic isn't changed, though a loop is required in the slow path to make sure all threads are aware of the end. Note that for now the soft-stop still takes time for group IDs > 1 as the poller is not yet started on these threads and needs to expire its timeout as there's no way to wake it up. But all threads are eventually stopped.	2022-07-01 19:15:15 +02:00
Willy Tarreau	03f9b35114	MEDIUM: tinfo: add a dynamic thread-group context The thread group info is not sufficient to represent a thread group's current state as it's read-only. We also need something comparable to the thread context to represent the aggregate state of the threads in that group. This patch introduces ha_tgroup_ctx[] and tg_ctx for this. It's indexed on the group id and must be cache-line aligned. The thread masks that were global and that do not need to remain global were moved there (want_rdv, harmless, idle). Given that all the masks placed there now become group-specific, the associated thread mask (tid_bit) now switches to the thread's local bit (ltid_bit). Both are the same for nbtgroups 1 but will differ for other values. There's also a tg_ctx pointer in the thread so that it can be reached from other threads.	2022-07-01 19:15:15 +02:00
Willy Tarreau	22b2a24eb2	CLEANUP: thread: remove thread_sync_release() and thread_sync_mask This function was added in 2.0 when reworking the thread isolation mechanism to make it more reliable. However it if fundamentally incompatible with the full isolation mechanism provided by thread_isolate_full() since that one will wait for all threads to become idle while the former will wait for all threads to finish waiting, causing a deadlock. Given that it's not used, let's just drop it entirely before it gets used by accident.	2022-07-01 19:15:15 +02:00
Willy Tarreau	cce203aae5	MINOR: thread: add a new all_tgroups_mask variable to know about active tgroups In order to kill all_threads_mask we'll need to have an equivalent for the thread groups. The all_tgroups_mask does just this, it keeps one bit set per enabled group.	2022-07-01 19:15:15 +02:00
Willy Tarreau	c6cf64bb5e	MINOR: thread: use ltid_bit in ha_tkillall() Since commit `cc7a11ee3` ("MINOR: threads: set the tid, ltid and their bit in thread_cfg") we ought not use (1UL << thr) to get the group mask for thread <thr>, but (ha_thread_info[thr].ltid_bit). ha_tkillall() needs this.	2022-07-01 19:15:15 +02:00
Willy Tarreau	1e7f0d68b0	MINOR: clock: use ltid_bit in clock_report_idle() Since commit `cc7a11ee3` ("MINOR: threads: set the tid, ltid and their bit in thread_cfg") we ought not use (1UL << thr) to get the group mask for thread <thr>, but (ha_thread_info[thr].ltid_bit). clock_report_idle() needs this. This also implies not using all_threads_mask anymore but taking the mask from the tgroup since it becomes relative now.	2022-07-01 19:15:15 +02:00
Willy Tarreau	adc1f52c92	MINOR: wdt: use ltid_bit in wdt_handler() Since commit `cc7a11ee3` ("MINOR: threads: set the tid, ltid and their bit in thread_cfg") we ought not use (1UL << thr) to get the group mask for thread <thr>, but (ha_thread_info[thr].ltid_bit). wdt_handler() needs this.	2022-07-01 19:15:14 +02:00
Willy Tarreau	38d0712748	MINOR: debug: use ltid_bit in ha_thread_dump() Since commit `cc7a11ee3` ("MINOR: threads: set the tid, ltid and their bit in thread_cfg") we ought not use (1UL << thr) to get the group mask for thread <thr>, but (ha_thread_info[thr].ltid_bit). ha_thread_dump() needs this.	2022-07-01 19:15:14 +02:00
Willy Tarreau	377e37a80f	MINOR: tinfo: add the mask of enabled threads in each group In order to replace the global "all_threads_mask" we'll need to have an equivalent per group. Take this opportunity for calling it threads_enabled and make sure which ones are counted there (in case in the future we allow to stop some).	2022-07-01 19:15:14 +02:00
Willy Tarreau	60fe4a95a2	MINOR: tinfo: replace the tgid with tgid_bit in tgroup_info Now that the tgid is accessible from the thread, it's pointless to have it in the group, and it was only set but never used. However we'll soon frequently need the mask corresponding to the group ID and the risk of getting it wrong with the +1 or to shift 1 instead of 1UL is important, so let's store the tgid_bit there.	2022-07-01 19:15:14 +02:00
Willy Tarreau	66ad98a772	MINOR: tinfo: add the tgid to the thread_info struct At several places we're dereferencing the thread group just to catch the group number, and this will become even more required once we start to use per-group contexts. Let's just add the tgid in the thread_info struct to make this easier.	2022-07-01 19:15:14 +02:00
Willy Tarreau	e7475c8e79	MEDIUM: tasks/fd: replace sleeping_thread_mask with a TH_FL_SLEEPING flag Every single place where sleeping_thread_mask was still used was to test or set a single thread. We can now add a per-thread flag to indicate a thread is sleeping, and remove this shared mask. The wake_thread() function now always performs an atomic fetch-and-or instead of a first load then an atomic OR. That's cleaner and more reliable. This is not easy to test, as broadcast FD events are rare. The good way to test for this is to run a very low rate-limited frontend with a listener that listens to the fewest possible threads (2), and to send it only 1 connection at a time. The listener will periodically pause and the wakeup task will sometimes wake up on a random thread and will call wake_thread(): frontend test bind :8888 maxconn 10 thread 1-2 rate-limit sessions 5 Alternately, disabling/enabling a frontend in loops via the CLI also broadcasts such events, but they're more difficult to observe since this is causing connection failures.	2022-07-01 19:15:14 +02:00
Willy Tarreau	dce4ad755f	MEDIUM: thread: add a new per-thread flag TH_FL_NOTIFIED to remember wakeups Right now when an inter-thread wakeup happens, we preliminary check if the thread was asleep, and if so we wake the poller up and remove its bit from the sleeping mask. That's not very clean since the sleeping mask cannot be entirely trusted since a thread that's about to wake up will already have its sleeping bit removed. This patch adds a new per-thread flag (TH_FL_NOTIFIED) to remember that a thread was notified to wake up. It's cleared before checking the task lists last, so that new wakeups can be considered again (since wake_thread() is only used to notify about task wakeups and FD polling changes). This way we do not need to modify a remote thread's sleeping mask anymore. As such wake_thread() now only tests and sets the TH_FL_NOTIFIED flag but doesn't clear sleeping anymore.	2022-07-01 19:15:14 +02:00
Willy Tarreau	555c192d14	MINOR: poller: update_fd_polling: wake a random other thread When enabling an FD that's only bound to another thread, instead of always picking the first one, let's pick a random one. This is rarely used (enabling a frontend, or session rate-limiting period ending), and has greater chances of avoiding that some obscure corner cases could degenerate into a poorly distributed load.	2022-07-01 19:15:14 +02:00
Willy Tarreau	962e5ba72b	MEDIUM: polling: make update_fd_polling() not care about sleeping threads Till now, update_fd_polling() used to check if all the target threads were sleeping, and only then would wake an owning thread up. This causes several problems among which the need for the sleeping_thread_mask and the fact that by the time we wake one thread up, it has changed. This commit changes this by leaving it to wake_thread() to perform this check on the selected thread, since wake_thread() is already capable of doing this now. Concretely speaking, for updt_fd_polling() it will mean performing one computation of an ffsl() before knowing the sleeping status on a global FD state change (which is very rare and not important here, as it basically happens after relaxing a rate-limit (i.e. once a second at beast) or after enabling a frontend from the CLI); thus we don't care.	2022-07-01 19:15:14 +02:00
Willy Tarreau	058b2c1015	MINOR: poller: centralize poll return handling When returning from the polling syscall, all pollers have a certain dance to follow, made of wall clock updates, thread harmless updates, idle time management and sleeping mask updates. Let's have a centralized function to deal with all of this boring stuff: fd_leaving_poll(), and make all the pollers use it.	2022-07-01 19:15:14 +02:00
Willy Tarreau	bdcd32598f	MINOR: thread: only use atomic ops to touch the flags The thread flags are touched a little bit by other threads, e.g. the STUCK flag may be set by other ones, and they're watched a little bit. As such we need to use atomic ops only to manipulate them. Most places were already using them, but here we generalize the practice. Only ha_thread_dump() does not change because it's run under isolation.	2022-07-01 19:15:14 +02:00
Willy Tarreau	f3efef4d60	MINOR: thread: make wake_thread() take care of the sleeping threads mask Almost every call place of wake_thread() checks for sleeping threads and clears the sleeping mask itself, while the function is solely used for there. Let's move the check and the clearing of the bit inside the function itself. Note that updt_fd_polling() still performs the check because its rules are a bit different.	2022-07-01 19:15:14 +02:00
Willy Tarreau	3fdacdddaf	MEDIUM: queue: revert to regular inter-task wakeups Now that the inter-task wakeups are cheap, there's no point in using task_instant_wakeup() anymore when dequeueing tasks. The use of the regular task_wakeup() is sufficient there and will preserve a better fairness: the test that went from 40k to 570k RPS now gives 580k RPS (down from 585k RPS with previous commit). This essentially reverts commit `27fab1dcb` ("MEDIUM: queue: use tasklet_instant_wakeup() to wake tasks").	2022-07-01 19:15:14 +02:00
Willy Tarreau	319d136ff9	MEDIUM: task: use regular eb32 trees for the run queues Since we don't mix tasks from different threads in the run queues anymore, we don't need to use the eb32sc_ trees and we can switch to the regular eb32 ones. This uses cheaper lookup and insert code, and a 16-thread test on the queues shows a performance increase from 570k RPS to 585k RPS.	2022-07-01 19:15:14 +02:00
Willy Tarreau	c958c70ec8	MINOR: task: replace global_tasks_mask with a check for tree's emptiness This bit field used to be a per-thread cache of the result of the last lookup of the presence of a task for each thread in the shared cache. Since we now know that each thread has its own shared cache, a test of emptiness is now sufficient to decide whether or not the shared tree has a task for the current thread. Let's just remove this mask.	2022-07-01 19:15:14 +02:00
Willy Tarreau	da195e8aab	MINOR: task: remove grq_total and use rq_total instead grq_total was only used to know how many tasks were being queued in the global runqueue for stats purposes, and that was transferred to the per thread rq_total counter once assigned. We don't need this anymore since we know where they are, so let's just directly update rq_total and drop that one.	2022-07-01 19:15:14 +02:00
Willy Tarreau	b17dd6cc19	MEDIUM: task: replace the global rq_lock with a per-rq one There's no point having a global rq_lock now that we have one shared RQ per thread, let's have one lock per runqueue instead.	2022-07-01 19:15:14 +02:00
Willy Tarreau	6f78038d72	MEDIUM: task: move the shared runqueue to one per thread Since we only use the shared runqueue to put tasks only assigned to known threads, let's move that runqueue to each of these threads. The goal will be to arrange an N*(N-1) mesh instead of a central contention point. The global_rqueue_ticks had to be dropped (for good) since we'll now use the per-thread rqueue_ticks counter for both trees. A few points to note: - the rq_lock stlil remains the global one for now so there should not be any gain in doing this, but should this trigger any regression, it is important to detect whether it's related to the lock or to the tree. - there's no more reason for using the scope-based version of the ebtree now, we could switch back to the regular eb32_tree. - it's worth checking if we still need TASK_GLOBAL (probably only to delete a task in one's own shared queue maybe).	2022-07-01 19:15:14 +02:00
Willy Tarreau	a4fb79b4a2	MINOR: task: make rqueue_ticks atomic The runqueue ticks counter is per-thread and wasn't initially meant to be shared. We'll soon have to share it so let's make it atomic. It's only updated when waking up a task, and no performance difference was observed. It was moved in the thread_ctx struct so that it doesn't pollute the local cache line when it's later updated by other threads.	2022-07-01 19:15:14 +02:00
Willy Tarreau	fc5de15baa	CLEANUP: task: remove the now unused TASK_GLOBAL flag TASK_GLOBAL was exclusively used by task_unlink_rq(), as such it can be dropped.	2022-07-01 19:15:14 +02:00
Willy Tarreau	eed3911a54	MINOR: task: replace task_set_affinity() with task_set_thread() The latter passes a thread ID instead of a mask, making the code simpler.	2022-07-01 19:15:14 +02:00
Willy Tarreau	159e3acf5d	MEDIUM: task: remove TASK_SHARED_WQ and only use t->tid TASK_SHARED_WQ was set upon task creation and never changed afterwards. Thus if a task was created to run anywhere (e.g. a check or a Lua task), all its timers would always pass through the shared timers queue with a lock. Now we know that tid<0 indicates a shared task, so we can use that to decide whether or not to use the shared queue. The task might be migrated using task_set_affinity() but it's always dequeued first so the check will still be valid. Not only this removes a flag that's difficult to keep synchronized with the thread ID, but it should significantly lower the load on systems with many checks. A quick test with 5000 servers and fast checks that were saturating the CPU shows that the check rate increased by 20% (hence the CPU usage dropped by 17%). It's worth noting that run_task_lists() almost no longer appears in perf top now.	2022-07-01 19:15:14 +02:00
Willy Tarreau	3b7a19c2a6	MINOR: applet: always use task_new_on() on applet creation Now that task_new_on() supports negative numbers, there's no need for checking the thread value nor falling back to task_new_anywhere().	2022-07-01 19:15:14 +02:00
Willy Tarreau	cb8542755e	MEDIUM: applet: only keep appctx_new_*() and drop appctx_new() This removes the mask-based variant so that from now on the low-level function becomes appctx_new_on() and it takes either a thread number or a negative value for "any thread". This way we can use task_new_on() and task_new_anywhere() instead of task_new() which will soon disappear.	2022-07-01 19:15:14 +02:00
Willy Tarreau	0ad00befc1	CLEANUP: task: remove thread_mask from the struct task It was not used anymore since everything moved to ->tid, so let's remove it.	2022-07-01 19:15:14 +02:00
Willy Tarreau	c44d08ebc4	MAJOR: task: replace t->thread_mask with 1<<t->tid when thread mask is needed At a few places where the task's thread mask. Now we know that it's always either one bit or all bits of all_threads_mask, so we can replace it with either 1<<tid or all_threads_mask depending on what's expected. It's worth noting that the global_tasks_mask is still set this way and that it's reaching its limits. Similarly, the task_new() API would deserve an update to stop using a thread mask and use a thread number instead. Similarly, task_set_affinity() should be updated to directly take a thread number. At this point the task's thread mask is not used anymore.	2022-07-01 19:15:14 +02:00
Willy Tarreau	29ffe26733	MAJOR: task: use t->tid instead of ffsl(t->thread_mask) to take the thread ID At several places we need to figure the ID of the first thread allowed to run a task. Till now this was performed using my_ffsl(t->thread_mask) but since we now have the thread ID stored into the task, let's use it instead. This is tagged major because it starts to assume that tid<0 is strictly equivalent to atleast2(thread_mask), and that as such, among the allowed threads are the current one.	2022-07-01 19:15:14 +02:00
Willy Tarreau	6ef52f4479	MEDIUM: task: add and preset a thread ID in the task struct The tasks currently rely on a mask but do not have an assigned thread ID, contrary to tasklets. However, in practice they're either running on a single thread or on any thread, so that it will be worth simplifying all this in order to ease the transition to the thread groups. This patch introduces a "tid" field in the task struct, that's either the number of the thread the task is attached to, or a negative value if the task is not bound to a thread, (i.e. its mask is all_threads_mask). The new ID is only set and updated but not used yet.	2022-07-01 19:15:14 +02:00
Willy Tarreau	8e5c53a6c9	MINOR: debug: remove mask support from "debug dev sched" The thread mask will not be used anymore, instead the thread id only is used. Interestingly it was already implemented in the parsing but not used. The single/multi thread argument is not needed anymore since it's sufficient to pass tid<0 to get a multi-threaded task/tasklet. This is in preparation for the removal of the thread_mask in tasks as only this debug code was using it!	2022-07-01 19:15:14 +02:00
Emeric Brun	7d392a592d	BUG/MEDIUM: ssl/fd: unexpected fd close using async engine Before 2.3, after an async crypto processing or on session close, the engine async file's descriptors were removed from the fdtab but not closed because it is the engine which has created the file descriptor, and it is responsible for closing it. In 2.3 the fd_remove() call was replaced by fd_stop_both() which stops the polling but does not remove the fd from the fdtab and the fd remains indefinitively in the fdtab. A simple replacement by fd_delete() is not a valid fix because fd_delete() removes the fd from the fdtab but also closes the fd. And the fd will be closed twice: by the haproxy's core and by the engine itself. Instead, let's set FD_DISOWN on the FD before calling fd_delete() which will take care of not closing it. This patch must be backported on branches >= 2.3, and it relies on this previous patch: MINOR: fd: add a new FD_DISOWN flag to prevent from closing a deleted FD As mentioned in the patch above, a different flag will be needed in 2.3.	2022-07-01 17:41:40 +02:00
Emeric Brun	f41a3f6762	MINOR: fd: add a new FD_DISOWN flag to prevent from closing a deleted FD Some FDs might be offered to some external code (external libraries) which will deal with them until they close them. As such we must not close them upon fd_delete() but we need to delete them anyway so that they do not appear anymore in the fdtab. This used to be handled by fd_remove() before 2.3 but we don't have this anymore. This patch introduces a new flag FD_DISOWN to let fd_delete() know that the core doesn't own the fd and it must not be closed upon removal from the fd_tab. This way it's totally unregistered from the poller but still open. This patch must be backported on branches >= 2.3 because it will be needed to fix a bug affecting SSL async. it should be adapted on 2.3 because state flags were stored in a different way (via bits in the structure).	2022-07-01 17:41:40 +02:00
Amaury Denoyelle	6befccd8a1	BUG/MINOR: mux-quic: do not signal FIN if gap in buffer Adjust FIN signal on Rx path for the application layer : ensure that the receive buffer has no gap. Without this extra condition, FIN was signalled as soon as the STREAM frame with FIN was received, even if we were still waiting to receive missing offsets. This bug could have lead to incomplete requests read from the application protocol. However, in practice this bug has very little chance to happen as the application layer ensures that the demuxed frame length is equivalent to the buffer data size. The only way to happen is if to receive the FIN STREAM as the H3 demuxer is still processing on a frame which is not the last one of the stream. This must be backported up to 2.6. The previous patch on ncbuf is required for the newly defined function ncb_is_fragmented(). MINOR: ncbuf: implement ncb_is_fragmented()	2022-07-01 15:55:32 +02:00
Amaury Denoyelle	e0a92a7e56	MINOR: ncbuf: implement ncb_is_fragmented() Implement a new status function for ncbuf. It allows to quickly report if a buffer contains data in a fragmented way, i.e. with gaps in between or at start of the buffer. To summarize, a buffer is considered as non-fragmented in the following cases : - a null or empty buffer - a full buffer - a buffer containing exactly one data block at the beginning, following by a gap until the end.	2022-07-01 15:54:23 +02:00
Amaury Denoyelle	36d4b5e31d	CLEANUP: mux-quic: adjust comment on qcs_consume() Since a previous refactoring, application protocol layer is not require anymore to call qcs_consume(). This function is now automatically used by the MUX itself.	2022-07-01 14:46:24 +02:00
Frédéric Lécaille	67fda16742	CLEANUP: h2: Typo fix in h2_unsubcribe() traces Very minor modification for the traces of this function.	2022-06-30 14:34:32 +02:00
Frédéric Lécaille	1b0707f3e7	MINOR: quic: Improvements for the datagrams receipt First we add a loop around recfrom() into the most low level I/O handler quic_sock_fd_iocb() to collect as most as possible datagrams before during its tasklet wakeup with a limit: we recvfrom() at most "maxpollevents" datagrams. Furthermore we add a local task list into the datagram handler quic_lstnr_dghdlr() which is passed to the first datagrams parser qc_lstnr_pkt_rcv(). This latter parser only identifies the connection associated to the datagrams then wakeup the highest level packet parser I/O handlers (quic_conn.*io_cb()) after it is done, thanks to the call to tasklet_wakeup_after() which replaces from now on the call to tasklet_wakeup(). This should reduce drastically the latency and the chances to fulfil the RX buffers at the QUIC connections level as reported in GH #1737 by Tritan. These modifications depend on this commit: "MINOR: task: Add tasklet_wakeup_after()" Must be backported to 2.6 with the previous commit.	2022-06-30 14:34:27 +02:00
Frédéric Lécaille	45a16295e3	MINOR: quic: Add new stats counter to diagnose RX buffer overrun Remove the call to qc_list_all_rx_pkts() which print messages on stderr during RX buffer overruns and add a new counter for the number of dropped packets because of such events. Must be backported to 2.6	2022-06-30 14:24:04 +02:00
Frédéric Lécaille	95a8dfb4c7	BUG/MINOR: quic: Dropped packets not counted (with RX buffers full) When the connection RX buffer is full, the received packets are dropped. Some of them were not taken into an account by the ->dropped_pkt counter. This very simple patch has no impact at all on the packet handling workflow. Must be backported to 2.6.	2022-06-30 14:24:04 +02:00
Frédéric Lécaille	ad548b54a7	MINOR: task: Add tasklet_wakeup_after() We want to be able to schedule a tasklet onto a thread after the current tasklet is done. What we have to do is to insert this tasklet at the head of the thread task list. Furthermore, we would like to serialize the tasklets. They must be run in the same order as the order in which they have been scheduled. This is implemented passing a list of tasklet as parameter (see <head> parameters) which must be reused for subsequent calls. _tasklet_wakeup_after_on() is implemented to accomplish this job. tasklet_wakeup_after_on() and tasklet_wake_after() are only wrapper macros around _tasklet_wakeup_after_on(). tasklet_wakeup_after_on() does exactly the same thing as _tasklet_wakeup_after_on() without having to pass the filename and line in the filename as parameters (usefull when DEBUG_TASK is enabled). tasklet_wakeup_after() hides also the usage of the thread parameter which is <tl> tasklet thread ID.	2022-06-30 14:24:04 +02:00
Amaury Denoyelle	a7a4c80ade	MINOR: qpack: properly handle invalid dynamic table references Return QPACK_DECOMPRESSION_FAILED error code when dealing with dynamic table references. This is justified as for now haproxy does not implement dynamic table support and advertizes a zero-sized table. The H3 calling function will thus reuse this code in a CONNECTION_CLOSE frame, in conformance with the QPACK RFC. This proper error management allows to remove obsolete ABORT_NOW guards.	2022-06-30 11:51:06 +02:00
Amaury Denoyelle	46e992d795	BUG/MINOR: qpack: abort on dynamic index field line decoding This is a complement to partial fix from commit `debaa04f9e` BUG/MINOR: qpack: abort on dynamic index field line decoding The main objective is to fix coverity report about usage of uninitialized variable when receiving dynamic table references. These references are invalid as for the moment haproxy advertizes a 0-sized dynamic table. An ABORT_NOW clause is present to catch this. A following patch will clean up this in order to properly handle QPACK errors with CONNECTION_CLOSE. This should fix github issue #1753. No need to backport as this was introduced in the current dev branch.	2022-06-30 11:51:06 +02:00
Amaury Denoyelle	2bc47863ec	MINOR: h3: handle errors on HEADERS parsing/QPACK decoding Emit a CONNECTION_CLOSE if HEADERS parsing function returns an error. This is useful to remove previous ABORT_NOW guards. For the moment, the whole connection is closed. In the future, it may be justified to only reset the faulting stream in some cases. This requires the implementation of RESET_STREAM emission.	2022-06-30 11:51:06 +02:00
Amaury Denoyelle	055de23b7d	BUG/MINOR: qpack: fix build with QPACK_DEBUG The local variable 't' was renamed 'static_tbl'. Fix its name in the qpack_debug_printf() statement which is activated only with QPACK_DEBUG mode. No need to backport as this was introduced in current dev branch.	2022-06-30 10:15:58 +02:00
Christopher Faulet	2b6777021d	MEDIUM: bwlim: Add support of bandwith limitation at the stream level This patch adds a filter to limit bandwith at the stream level. Several filters can be defined. A filter may limit incoming data (upload) or outgoing data (download). The limit can be defined per-stream or shared via a stick-table. For a given stream, the bandwith limitation filters can be enabled using the "set-bandwidth-limit" action. A bandwith limitation filter can be used indifferently for HTTP or TCP stream. For HTTP stream, only the payload transfer is limited. The filter is pretty simple for now. But it was designed to be extensible. The current design tries, as far as possible, to never exceed the limit. There is no burst.	2022-06-24 14:06:26 +02:00
Frédéric Lécaille	628e89cfae	BUILD: quic+h3: 32-bit compilation errors fixes In GH #1760 (which is marked as being a feature), there were compilation errors on MacOS which could be reproduced in Linux when building 32-bit code (-m32 gcc option). Most of them were due to variables types mixing in QUIC_MIN macro or using size_t type in place of uint64_t type. Must be backported to 2.6.	2022-06-24 12:13:53 +02:00
Fr�d�ric L�caille	2bed1f166e	BUG/MAJOR: quic: Big RX dgrams leak with POST requests This previous commit: "BUG/MAJOR: Big RX dgrams leak when fulfilling a buffer" partially fixed an RX dgram memleak. There is a missing break in the loop which looks for the first datagram attached to an RX buffer dgrams list which may be reused (because consumed by the connection thread). So when several dgrams were consumed by the connection thread and are present in the RX buffer list, some are leaked because never reused for ever. They are removed for their list. Furthermore, as commented in this patch, there is always at least one dgram object attached to an RX dgrams list, excepted the first time we enter this I/O handler function for this RX buffer. So, there is no need to use a loop to lookup and reuse the first datagram in an RX buffer dgrams list. This isssue was reproduced with quiche client with plenty of POST requests (100000 streams): cargo run --bin quiche-client -- https://127.0.0.1:8080/helloworld.html --no-verify -n 100000 --method POST --body /var/www/html/helloworld.html and could be reproduce with GET request. This bug was reported by Tristan in GH #1749. Must be backported to 2.6.	2022-06-23 21:57:09 +02:00
Fr�d�ric L�caille	19ef6369b5	BUG/MAJOR: quic: Big RX dgrams leak when fulfilling a buffer When entering quic_sock_fd_iocb() I/O handler which is responsible of recvfrom() datagrams, the first thing which is done it to try to reuse a dgram object containing metadata about the received datagrams which has been consumed by the connection thread. If this object could not be used for any reason, so when we "goto out" of this function, we must release the memory allocated for this objet, if not it will leak. Most of the time, this happened when we fulfilled a buffer as reported in GH #1749 by Tristan. This is why we added a pool_free() call just before the out label. We mark <new_dgram> as NULL when it successfully could be used. Thank you for Tristan and Willy for their participation on this issue. Must be backported to 2.6.	2022-06-23 20:40:01 +02:00
Fr�d�ric L�caille	0c535683ee	BUG/MINOR: quic: Wrong reuse of fulfilled dgram RX buffer After having fulfilled a buffer, then marked it as full, we must consume the remaining space. But to do that, and not to erase the already existing data, we must check there is not remaining data in after the tail of the buffer (between the tail and the head). This is done adding a condition to test that adding the number of bytes from the remaining contiguous space to the tail does not pass the wrapping postion in the buffer. Must be backported to 2.6.	2022-06-23 20:39:19 +02:00
Willy Tarreau	27061cd144	MEDIUM: debug: improve DEBUG_MEM_STATS to also report pool alloc/free Sometimes using "debug dev memstats" can be frustrating because all pool allocations are reported through pool-os.h and that's all. But in practice there's nothing wrong with also intercepting pool_alloc, pool_free and pool_zalloc and report their call counts and locations, so that's what this patch does. It only uses an alternate set of macroes for these 3 calls when DEBUG_MEM_STATS is defined. The outputs are reported as P_ALLOC (for both pool_malloc() and pool_zalloc()) and P_FREE (for pool_free()).	2022-06-23 11:58:01 +02:00
Willy Tarreau	b8dec4a01a	CLEANUP: pool/tree-wide: remove suffix "_pool" from certain pool names A curious practise seems to have started long ago and contaminated various code areas, consisting in appending "_pool" at the end of the name of a given pool. That makes no sense as the name is only used to name the pool in diags such as "show pools", and since names are truncated there, this adds some confusion when analysing the dump outputs. Let's just clean all of them at once. there were essentially in SSL and QUIC.	2022-06-23 11:49:09 +02:00
Willy Tarreau	47af317389	BUG/MINOR: stream: only free the req/res captures when set There's a subtle bug in stream_free() when releasing captures. The pools may be NULL when no capture is defined, and the calls to pool_free() are inconditional. The only reason why this doesn't cause trouble is because the pointer to be freed is always NULL in this case and we don't go further down the chain. That's particularly ugly and it complicates debugging, so let's only call these ones when the pointers are set. There's no impact on running code, it only fools those trying to debug pools manually. There's no need to backport it though it unless it helps for debugging sessions.	2022-06-23 11:49:09 +02:00
Christopher Faulet	aa55640b8c	MINOR: freq_ctr: Add a function to get events excess over the current period freq_ctr_overshoot_period() function may be used to retrieve the excess of events over the current period for a givent frequency counter, ignoring the history. It is a way compare the "current rate" (the number of events over the current period) to a given rate and estimate the excess of events. It may be used to safely add new events, especially at the begining of the current period for a frequency counter with large period. This way, it is possible to smoothly add events during the whole period without quickly consuming all the quota at the beginning of the period and waiting for the next one to be able to add new events.	2022-06-22 18:33:27 +02:00
Christopher Faulet	dbbdb25f1c	BUG/MINOR: http-fetch: Use integer value when possible in "method" sample fetch Because of the previous fix, if the HTTP parsing is performed when the "method" sample fetch is called, we always rely on the string representation of the request method. Indeed, if no parsing was performed when the "method" sample fetch is called, the transaction method is HTTP_METH_OTHER because it was just initialized. However, without this patch, in this case, we always retrieve the method by reading the request start-line. Now, when the method is HTTP_METH_OTHER, we systematically try to parse the request but the method is tested once again after the parsing to be able to use the integer representation when possible. This patch must be backported as far as 2.0.	2022-06-22 17:50:54 +02:00
Christopher Faulet	5eb67f5d74	BUG/MINOR: http-ana: Set method to HTTP_METH_OTHER when an HTTP txn is created This patch is required to fix "method" sample fetch. But it make sense to initialize the method of an HTTP transaction to HTTP_METH_OTHER. This way, before the request parsing, the method is considered as unknown except if we are able to retrieve the request start-line. It is especially important for TCP streams. About the "method" sample fetch, this patch is a way to be sure no random method is returned when the sample fetch is used on a TCP stream before any HTTP parsing. This patch must be backported as far as 2.0.	2022-06-22 17:50:54 +02:00
Frédéric Lécaille	77ac6f5667	BUG/MINOR: quic: Missing acknowledgments for trailing packets This bug was revealed by key_update QUIC tracker test. During this test, after the handshake has succeeded, the client sends a 1-RTT packet containing only a PING frame. On our side, we do not acknowledge it, because we have no ack-eliciting packet to send. This is not correct. We must acknowledge all the ack-eliciting packets unconditionally. But we must not send too much ACK frames: every two packets since we have sent an ACK frame. This is the test (nb_aepkts_since_last_ack >= QUIC_MAX_RX_AEPKTS_SINCE_LAST_ACK) which ensure this is the case. But this condition must be checked at the very last time, when we are building a packet and when an acknowledgment is required. This is not to qc_may_build_pkt() to do that. This boolean function decides if we have packets to send. It must return "true" if there is no more ack-eliciting packets to send and if an acknowledgement is required. We must also add a "force_ack" parameter to qc_build_pkt() to force the acknowledments during the handshakes (for each packet). If not, they will be sent every two packets. This parameter must be passed to qc_do_build_pkt() and be checked alongside the others conditions when building a packet to decide to add an ACK frame or not to this packet. Must be backported to 2.6.	2022-06-22 15:47:52 +02:00
Remi Tricot-Le Breton	1bad7db4a1	BUG/MINOR: ssl: Do not look for key in extra files if already in pem A bug was introduced by commit `9bf3a1f67e` "BUG/MINOR: ssl: Fix crash when no private key is found in pem". If a private key is already contained in a pem file, we will still look for a .key file and load its private key if it exists when we should not. This patch should be backported to all branches where the original fix was backported (all the way to 2.2).	2022-06-22 10:45:47 +02:00
Willy Tarreau	d543ae0e68	BUILD: ssl_ckch: fix "maybe-uninitialized" build error on gcc-9.4 + ARM As reported in issue #1755, gcc-9.3 and 9.4 emit a "maybe-uninitialized" warning in cli_io_handler_commit_cafile_crlfile() because it sees that when the "path" variable is not set, we're jumping to the error label inside the loop but cannot see that the new state will avoid the places where the value is used. Thus it's a false positive but a difficult one. Let's just preset the value to NULL to make it happy. This was introduced in 2.7-dev by commit `ddc8e1cf8` ("MINOR: ssl_ckch: Simplify I/O handler to commit changes on CA/CRL entry"), thus no backport is needed for now.	2022-06-22 05:45:34 +02:00
Willy Tarreau	c7a8a3c7bd	MINOR: intops: add a function to return a valid bit position from a mask Sometimes we need to be able to signal one thread among a mask, without caring much about which bit will be picked. At the moment we use ffsl() for this but this sometimes results in imbalance at certain specific places where the same first thread in a set is always the same one that is selected. Another approach would consist in using the rank finding function but it requires a popcount and a setup phase, and possibly a modulo operation depending on the popcount, which starts to be very expensive. Here we take a different approach. The idea is an input bit value is passed, from 0 to LONGBITS-1, and that as much as possible we try to pick the bit matching it if it is set. Otherwise we look at a mirror position based on a decreasing power of two, and jump to the side that still has bits left. In 6 iterations it ends up spotting one bit among 64 and the operations are very cheap and optimizable. This method has the benefit that we don't care where the holes are located in the mask, thus it shows a good distribution of output bits based on the input ones. A long-time test shows an average of 16 cycles, or ~4ns per lookup at 3.8 GHz, which is about twice as fast as using the rank finding function. Just like for that one, the code was stored into tools.c since we don't have a C file for intops.	2022-06-21 20:29:57 +02:00
William Lallemand	0a012aa16b	BUG/MEDIUM: mworker: use default maxconn in wait mode In bug #1751, it was reported that haproxy is consumming too much memory since the 2.4 version. This is because of a change in the master, which loses completely its configuration in wait mode, and lose its maxconn. Without the maxconn, haproxy will try to compute one itself, and will allocate RAM consequently, too much in our case. Which means the master will have a too high maxconn and too much RAM allocated. The patch fixes the issue by setting the maxconn to the default value when re-executing the master in wait mode. Must be backported as far as 2.5.	2022-06-21 14:22:49 +02:00
Frédéric Lécaille	4f5777a415	MINOR: quic: Dump version_information transport parameter Implement quic_tp_version_info_dump() to dump such a transport parameter (only remote). Call it from quic_transport_params_dump() which dump all the transport parameters. Can be backported to 2.6 as it's useful for debugging.	2022-06-21 11:07:39 +02:00
Frédéric Lécaille	57bddbcbbb	BUG/MINOR: quic: Acknowledgement must be forced during handshake All packets received during hanshakes must be acknowledged asap. This was not the case for Handshake packets received. At this time, this had no impact because the client has often only one Handshake packet to send and last handshake to be sent on our side always embeds an HANDSHAKE_DONE frame which leads the client to consider it has no more handshake packet to send. Add <force_ack> to qc_may_build_pkt() to force an ACK frame to be sent. Set this parameter to 1 when sending packets from Initial or Handshake packet number spaces, 0 when sending only Application level packet. Must be backported to 2.6.	2022-06-21 11:00:16 +02:00
William Lallemand	cb6c5f4683	BUG/MEDIUM: ssl/cli: crash when crt inserted into a crt-list The crash occures when the same certificate which is used on both a server line and a bind line is inserted in a crt-list over the CLI. This is quite uncommon as using the same file for a client and a server certificate does not make sense in a lot of environments. This patch fixes the issue by skipping the insertion of the SNI when no bind_conf is available in the ckch_inst. Change the reg-test to reproduce this corner case. Should fix issue #1748. Must be backported as far as 2.2. (it was previously in ssl_sock.c)	2022-06-20 17:27:49 +02:00
Amaury Denoyelle	debaa04f9e	BUG/MINOR: qpack: abort on dynamic index field line decoding Add an ABORT_NOW() clause if indexed field line referred to the dynamic table. This is required as current haproxy QPACK implementation does not support the dynamic table. Note that this should not happen as haproxy explicitely advertizes a null-sized dynamic table to the other peer. This shoud fix github issue #1753. No need to backport as this was introduced by commit `b666c6b26e` MINOR: qpack: improve decoding function	2022-06-20 15:56:01 +02:00
Amaury Denoyelle	23f908ccd6	BUG/MINOR: quic: free rejected Rx packets Free Rx packet in the datagram handler if the packet was not taken in charge by a quic_conn instance. This is reflected by the packet refcount which is null. A packet can be rejected for a variety of reasons. For example, failed decryption, no Initial token and Retry emission or for datagram null padding. This patch should resolve the Rx packets memory leak observed via "show pools" with the previous commit `2c31e12936` BUG/MINOR: quic: purge conn Rx packet list on release This specific memory leak instance was reproduced using quiche client which uses null datagram padding. This should partially resolve github issue #1751. It must be backported up to 2.6.	2022-06-20 15:04:07 +02:00
Amaury Denoyelle	2c31e12936	BUG/MINOR: quic: purge conn Rx packet list on release When releasing a quic_conn instance, free all remaining Rx packets in quic_conn.rx.pkt_list. This partially fixes a memory leak on Rx packets which can be observed after several QUIC connections establishment. This should partially resolve github issue #1751. It must be backported up to 2.6.	2022-06-20 14:59:52 +02:00
Frédéric Lécaille	483499dc22	BUG/MINOR: quic_stats: Duplicate "quic_streams_data_blocked_bidi" field name As reported by broxio in GH #1757, there was a duplication field name for "quic_streams_data_blocked_bidi", due to a copy and paste without renaming I guess. Must be backported to 2.6.	2022-06-20 14:57:19 +02:00
Frédéric Lécaille	2aebaa49b1	BUG/MINOR: quic: Unexpected half open connection counter wrapping This counter must be incremented only one time by connection and decremented as soon as the handshake has failed or succeeded. This is a gauge. Under certain conditions this counter could be decremented twice. For instance after having received a TLS alert, then upon SSL_do_handshake() failure. To stop having to deal to all the current combinations which can lead to such a situation (and the next to come), add a connection flag to denote if this counter has been already decremented for a connection. So, this counter must be decremented only if this flag has not been already set. Must be backported up to 2.6.	2022-06-20 14:57:09 +02:00
Frédéric Lécaille	b1cb958581	BUILD: quic: Wrong HKDF label constant variable initializations Non constant expressions were used to initialize constant variables leading to such compilation errors: src/xprt_quic.c:66:3: error: initializer element is not a constant expression .key_label_len = strlen(QUIC_HKDF_KEY_LABEL_V1), Reproduced with CC=gcc-4.9 compilation option. Fix using macros for each HKDF label.	2022-06-20 14:50:19 +02:00
Willy Tarreau	177aed56dc	MEDIUM: debug: detect redefinition of symbols upon dlopen() In order to better detect the danger caused by extra shared libraries which replace some symbols, upon dlopen() we now compare a few critical symbols such as malloc(), free(), and some OpenSSL symbols, to see if the loaded library comes with its own version. If this happens, a warning is emitted and TAINTED_REDEFINITION is set. This is important because some external libs might be linked against different libraries than the ones haproxy was linked with, and most often this will end very badly (e.g. an OpenSSL object is allocated by haproxy and freed by such libs). Since the main source of dlopen() calls is the Lua lib, via a "require" statement, it's worth trying to show a Lua call trace when detecting a symbol redefinition during dlopen(). As such we emit a Lua backtrace if Lua is detected as being in use.	2022-06-19 17:58:32 +02:00
Willy Tarreau	40dde2d5c1	MEDIUM: debug: add a tainted flag when a shared library is loaded Several bug reports were caused by shared libraries being loaded by other libraries or some Lua code. Such libraries could define alternate symbols or include dependencies to alternate versions of a library used by haproxy, making it very hard to understand backtraces and analyze the issue. Let's intercept dlopen() and set a new TAINTED_SHARED_LIBS flag when it succeeds, so that "show info" makes it visible that some external libs were added. The redefinition is based on the definition of RTLD_DEFAULT or RTLD_NEXT that were previously used to detect that dlsym() is usable, since we need it as well. This should be sufficient to catch most situations.	2022-06-19 17:58:32 +02:00
Willy Tarreau	0b7b639d7e	MINOR: hlua: add a new hlua_show_current_location() function This function may be used to try to show where some Lua code is currently being executed. It tries hard to detect the initialization phase, both for the global and the per-thread states, and for runtime states. This intends to be used by error handlers to provide the users with indications about what Lua code was being executed when the error triggered.	2022-06-19 17:58:32 +02:00
Willy Tarreau	5c143404ea	MINOR: hlua: don't dump empty entries in hlua_traceback() Calling hlua_traceback() sometimes reports empty entries looking like: [C]: ? These ones correspond to certain internal C functions of the Lua library, but they do not provide any information and even complicate the interpretation of the dump. Better just skip them.	2022-06-19 17:58:32 +02:00
Christopher Faulet	a892b7f15f	BUG/MINOR: log: Properly test connection retries to fix dontlog-normal option The commit `731c8e6cf` ("MINOR: stream: Simplify retries counter calculation") introduced a regression. It broke the dontlog-normal option because the test on the connection retries counter was not updated accordingly. This patch should fix the issue #1754. It must be backported to 2.6.	2022-06-17 14:53:21 +02:00
Christopher Faulet	82af3c684d	CLEANUP: stconn: Don't expect to have no sedesc on detach The stream connector must always have a defined sedesc. So there is no reason to test it when the stconn is detached from the endpoint.	2022-06-17 13:25:02 +02:00
Christopher Faulet	9b8d7a11c0	MINOR: stream: Rely on stconn flags to abort stream destructive upgrade On destructive connection upgrade, instead of using the new mux name to abort the old stream, we can relay on the stream connector flags. If it is detached after the upgrade, it means the stream will not be resused by the new mux and it must be aborted. This patch may be backported to 2.6.	2022-06-17 13:25:02 +02:00
Christopher Faulet	b68f77d626	BUG/MEDIUM: stream: Properly handle destructive client connection upgrades When the protocol is changed for a client connection at the stream level (from TCP to H1/H2), there are two cases. The stream may be reused or not. The first case, when the stream is reused is working. The second one is buggy since the conn-stream refactoring and leads to a crash. In this case, the new mux don't reuse the stream. It must be silently aborted. However, its front stream connector is still referencing the connection. So it must be detached. But it must be performed in two stages, to be sure to not loose the context for the upgrade and to be able to rollback on error. So now, before the upgrade, we prepare to detach the stconn and it is finally detached if the upgrade succeeds. There is a trick here. Because we pretend the stconn is detached but its state is preserved. This patch must be backported to 2.6.	2022-06-17 13:25:02 +02:00
Willy Tarreau	9b3aa63df7	BUG/MINOR: task: fix thread assignment in tasklet_kill() tasklet_kill() was introduced in 2.5-dev4 with commit `7b368339a` ("MEDIUM: task: implement tasklet kill"), but a comparison error there makes tasklets killed on thread 1 assigned to the killing thread. Fortunately, the function was finally not used so there's no harm right now, hence the minor tag, but this must be fixed and backported in case a later fix relies on it. This should be backported to 2.5.	2022-06-16 18:17:44 +02:00
Frédéric Lécaille	e06f7459fa	CLEANUP: quic: Remove any reference to boringssl I do not think we will support boringssl for QUIC soon ;)	2022-06-16 15:58:48 +02:00
Frédéric Lécaille	301425b880	MEDIUM: quic: Compatible version negotiation implementation (draft-08) At this time haproxy supported only incompatible version negotiation feature which consists in sending a Version Negotiation packet after having received a long packet without compatible value in its version field. This version value is the version use to build the current packet. This patch does not modify this behavior. This patch adds the support for compatible version negotiation feature which allows endpoints to negotiate during the first flight or packets sent by the client the QUIC version to use for the connection (or after the first flight). This is done thanks to "version_information" parameter sent by both endpoints. To be short, the client offers a list of supported versions by preference order. The server (or haproxy listener) chooses the first version it also supported as negotiated version. This implementation has an impact on the tranport parameters handling (in both direcetions). Indeed, the server must sent its version information, but only after received and parsed the client transport parameters). So we cannot encode these parameters at the same time we instantiated a new connection. Add QUIC_TP_DRAFT_VERSION_INFORMATION(0xff73db) new transport parameter. Add tp_version_information new C struct to handle this new parameter. Implement quic_transport_param_enc_version_info() (resp. quic_transport_param_dec_version_info()) to encode (resp. decode) this parameter. Add qc_conn_finalize() which encodes the transport parameters and configure the TLS stack to send them. Add ->negotiated_ictx quic_conn C struct new member to store the Initial QUIC TLS context for the negotiated version. The Initial secrets derivation is version dependent. Rename ->version to ->original_version and add ->negotiated_version to this C struct to reflect the QUIC-VN RFC denomination. Modify most of the QUIC TLS API functions to pass a version as parameter. Export the QUIC version definitions to be reused at least from quic_tp.c (transport parameters. Move the token check after the QUIC connection lookup. As this is the original version which is sent into a Retry packet, and because this original version is stored into the connection, we must check the token after having retreived this connection. Add packet version to traces. See https://datatracker.ietf.org/doc/html/draft-ietf-quic-version-negotiation-08 for more information about this new feature.	2022-06-16 15:58:48 +02:00
Frédéric Lécaille	e17bf77218	MINOR: quic: Released QUIC TLS extension for QUIC v2 draft This is not clear at all how to distinguish a QUIC draft version number from a released one. And among these QUIC draft versions, which one must use the draft QUIC TLS extension. According to the QUIC implementations which support v2 draft, the TLS extension (transport parameters) to be used is the released one (TLS_EXTENSION_QUIC_TRANSPORT_PARAMETERS). As the unique QUIC draft version we support is 0xff00001d and as at this time the unique version with 0xff as most significant byte is this latter which must use the draft TLS extension, we select the draft TLS extension (TLS_EXTENSION_QUIC_TRANSPORT_PARAMETERS_DRAFT) only for such versions with 0xff as most signification byte.	2022-06-16 14:56:24 +02:00
Frédéric Lécaille	86845c5171	MEDIUM: quic: Add QUIC v2 draft support This is becoming difficult to handle the QUIC TLS related definitions which arrive with a QUIC version (draft or not). So, here we add quic_version C struct which does not define only the QUIC version number, but also the QUIC TLS definitions which depend on a QUIC version. Modify consequently all the QUIC TLS API to reuse these definitions through new quic_version C struct. Implement quic_pkt_type() function which return a packet type (0 up to 3) depending on the QUIC version number. Stop harding the Retry packet first byte in send_retry(): this is not more possible because the packet type field depends on the QUIC version. Also modify quic_build_packet_long_header() for the same reason: the packet type depends on the QUIC version. Add a quic_version C struct member to quic_conn C struct. Modify qc_lstnr_pkt_rcv() to set this member asap. Remove the version member from quic_rx_packet C struct: a packet is attached asap to a connection (or dropped) which is the unique object which should store the QUIC version. Modify qc_pkt_is_supported_version() to return a supported quic_version C struct from a version number. Add Initial salt for QUIC v2 draft (initial_salt_v2_draft).	2022-06-16 14:56:24 +02:00

... 19 20 21 22 23 ...

16085 Commits