Fixes #35416, by:
- First, a general CI cleanup to use one network per CI container and remove it after success
- Then, to use one subnet per container and a unique IP in that subnet (to avoid collisions)
The networks were left around for the next run, which may cause issues
when they are changed in a future CI change.
Fix it by using one network per container, and by cleaning it up after a
successful run.
This also allows to harden the script by removing the check=False on
creation.
On a failed run, networks will remain dangling and need manual pruning
(same behaviour as before).
This avoids collisions between concurrently running containerized CI jobs.
<!--e57a25ab6845829454e8d69fc972939a-->
The following sections might be updated with supplementary metadata relevant to reviewers and maintainers.
<!--006a51241073e994b41acfe9ec718e94-->
<!--021abf342d371248e50ceaed478a90ca-->
See the guideline and AI policy for information on the review process.
| Type | Reviewers |
|---|---|
| ACK | willcl-ark, sedited |
If your review is incorrectly listed, please copy-paste <code><!--meta-tag:bot-skip--></code> into the comment that the bot should ignore.
<!--5faf32d7da4f0f540f40219e4f7537a3-->
@willcl-ark This is a slightly different approach, based on your attempt from May (with my feedback addressed from #35417#pullrequestreview-4410066861 without causing a ci failure :sweat_smile: )
You may be qualified to nack/ack. Also, do you want to be co-author of any of the commits?
Yes was just taking a look here.
do you want to be co-author of any of the commits?
no thanks, I'm fine :)
I seem to recall wondering about this previously, but do you think we should try to remove the network before creating, to avoid bumping into network creation issues after a failed run (where the user presumably kills the container but will forget to rm the network)?
I don't mind strongly, but perhaps with a dedicated network per container this is even "safer" than before?
I don't think the network can be removed before creating: The only case where it exists is when the run failed (and the container still exists). A network can not be deleted when the container exists, IIRC? Maybe there could be a CI_FORCE_FRESH_CONTAINER=1 setting?
<!--
Maybe there could be a CI_FORCE_FRESH_CONTAINER=1 setting?
I don't think this is worth it, the error message is fine.
ACK fa966fd76764a0e150f654d7c2760e58226edba5
Code looks fine.
Manually checked 3 or 4 Ci runs too to ensure they were using the correct names, e.g. on netbsd:
+ docker network create --ipv6 --subnet=1111:1111:14::/112 --subnet=11.11.20.0/24 ci_netbsd_cross-net
ACK fa966fd76764a0e150f654d7c2760e58226edba5
utACK fa966fd76764a0e150f654d7c2760e58226edba5