New Pip resolver takes a long time to complete(github.com) |
New Pip resolver takes a long time to complete(github.com) |
They made a half-assed attempt when pypa/pip added the "data-requires-python" tag, but that covers python interpreter version only, it needed to have been done for all dependencies.
What's especially irksome is that Debian and RPM both solved this 20+ years ago and python has refused to learn any lessons.
I wish different language/platform communities were better at learning from each other (and this does go in all directions). In the field of software these days, we don't do much learning from prior art. It's just too hard to keep up with it all.
Heck, I recall this happening with rubygems -- the introductio of the dependencies API to efficiently get the minimum data needing for resolving dependencies before downloading actual packages, and then several iterations on it -- and I can't find any actual documentary evidence of it to share right now. I don't know how anyone WOULD use it as prior art; the source code for a fairly complex project in a language you aren't familiar with isn't going to work, even if you knew to go look for it, which why would you.
See the “requires’ key here: https://pypi.org/pypi/Django/json
The TL;DR is that it probably requires a PEP.
FD: I've been working on adjacent features in PyPI.
Then realised it would have this problem a few days before I'd planned to release, and didn't ship it.
I am feeling very lucky right now.
Are there other languages that have dependencies decided at installtime?
setup(
install_requires=[random.choice(["urllib3", "requests"])]
)
This example wouldn't make any sense, but you could imagine installing different dependencies for x86 CPUs or something via a runtime check, and there are lots of packages that use this for checking python versions even though there's now a static way to do that.So that leads to this situation, from another comment:
"And as an example, [the botocore package, which has releases nearly daily] depends on python-dateutil>=2.1,<3.0.0. So if [your dependency constraints are] to install python-dateutil 3.0.0 and botocore, pip will have to backtrack through every release of botocore before it can be sure that there isn't one that works with dateutil 3.0.0. ... And worse still, if an ancient version of botocore does have an unconstrained dependency on python-dateutil, we could end up installing it with dateutil 3.0.0, and have a system that, while technically consistent, doesn't actually work."
Sounds like there's a long term plan that could fix this situation. Binary wheels already have the needed metadata in a way that could be exposed by PyPI via a fast "fetch dependency constraints for all versions" API, but isn't yet. And for source dists there's a very new plan ( https://www.python.org/dev/peps/pep-0643/ ) to let them indicate that they don't modify install_requires at runtime, so their deps could also be exposed via API, but the ecosystem will have to catch up with that.
I dunno what pip does in the meantime, though!
(I'm a python dev but haven't followed this beyond skimming the bug, so I hope I'm getting this right.)
It's NP-complete in general I think, definitely the case in Haskell-land. I think OCaml sets an explicit timeout for dependency resolution?
> the dependency graph cannot be computed without installing many versions of all dependencies?
Does PyPi not have an index or something? (with statically known package bounds?)
The pip section in a env file is just a list of arguments passed through to the pip install command. Prior to pip 20.3 we had to add `--use-feature=2020-resolver` to get an install that resolved for our teams that used mamba.
You can install a pip package inside a conda environment. But when running `mamba install` or `conda install`, pip is not involved at all.
Since using GNU Guix though, I'm so glad it doesn't have a dependency resolver as part of building or installing packages! It's so much better for it, no slow or unpredictable resolving, you know what it's going to do.
I think this is one reason why I've never used pip for managing Python software, I've only ever used Debian, and then Guix.
If I'm updating a library that has a bunch of dependent packages, it's hard to know whether or not something will break downstream. Sometimes this is unavoidable and you really do need to just test everything thoroughly, but sometimes the dependent packages have more knowledge of what library versions they need... and Guix doesn't seem to be aware of this.
What we need is to be able to use a resolver such as included by pip or poetry, to build up our package set. In Nixpkgs this is nowadays unfortunately done manually, and I suppose the same goes for Guix. In Nixpkgs the reason is simple: too eager pinning makes it impossible to resolve a package set that works with the entire set.
Now that pip has a resolver what is needed is a way to use constraints not to set only lower and upper bounds, but to enforce a version when resolving. That makes it usable for downstream integrators to construct their primary package set. One could then even make the next step and construct "stable" sets that extend the primary set.
Guix package definitions are truely code though, so if you want to generate packages on the fly by using a dependency resolver, you can totally write some code to make that happen.
With respect to inefficiency, what do you mean? It's quite time efficient when building and installing to not have to attempt to resolve dependencies.
Conda installs conda packages and conda uses pip to install pip packages. However a pip package can be converted to a conda package, and then in that case the dependency will be installed by conda and not pip.
You can install a pip package in a conda env. But this is actually not recommended.
When using `mamba install` or `conda install`, pip is not involved at all.
Ouch, what languages are you using where this is true? For example, JS / NPM / Yarn are absolutely blowing pip out of the water. I guess I can imagine Java or C++ users having your perspective though
There are a lot of smart people working on these things in each ecosystem, and when you think some package manager is far surperior over others in every way, you are more than often simply wrong. Or saying it the other way (and paraphrasing a commenter from another thread), the only package manager you think is good is from the ecosystem you are not deeply familiar with.
Or maybe, the only package manager you think is good is the one that has features you value and doesn't fail in ways you wouldn't expect?
For example, I have been trying to work with numpy and pandas a bit - two of the biggest Python libraries - and use them on my MacBook (itself a popular item). These installs fail, in the middle of an ungrokable stack of install logs. I have to shuffle thru useless messages to eventually track down the source, and then try to find a new compatible version. So sure, maybe there's native code in there, but I think claiming "useful summaries are not the package manager's job" is silly
But I also have had other terrible experiences: packaging is a bit overwrought; the terrible import/export/module system in Python means your dependency names have nothing to do with where you import from; SSL Certificate errors crop up at random on domains with valid certificate.
I am sure it's good enough to use, but I think it's a bit pie-in-the-sky to claim all package managers are good and you just have to be in the community. Shitty software exists
It's going to burn disk space like crazy, isn't it? If I install packages foo-1.0 and bar-1.0 and foo-1.0 uses glibc-2.31 and bar-1.0 uses glibc-2.30 then I now have 2 versions of glibc... but that scales to every package I install and every library every one of them uses. Of course, this can be fixed... by automatically building new versions of every package any time any of their dependencies changes, in which case we're probably not wasting much disk because we traded and are now burning CPU time like there's no tomorrow (and network, and disk I/O, and memory, and anything else used in package builds). Basically, this sounds like reinventing static binaries, with all the downsides thereof.
Because of the immutable store, you can do file level deduplication, so if you have multiple versions of the same packages, you can deduplicate the identical files.
I think the worries about disk usage are relevant, but only on systems with small amounts of storage. These are still relevant though, and it's an important area to improve on. As for burning CPU time, I don't think there's a perfect solution to avoid this, but I think Guix is pretty good. Guix provides substitutes, so you don't have to build things locally on every machine (I'm looking at you Rubygems, pip and Python stuff is pretty bad also).
It's disastrous for security patches, only highly inconvenient for things like performance improvement releases. But this is why we have dependency resolving, right? What am I missing?
There is an issue here of rebuilding all those dependent packages with the updated A, especially if it's something like glibc. Guix includes a mechanism called grafts that allows for package replacements, which allows avoiding this, and this is often used for releasing security fixes.
> you're right in that normally package definitions specify the exact dependencies it the code (they're hardcoded).
I thought that meant that package X might specify a dependency on A version 2.4.2 exactly. So if you want it to use A 2.4.3 instead, a new release of package X would have to be created, specifying 2.4.3.
But I think this is not in fact what you mean? In which case I don't yet understand the system you are describing and what you mean by not doing dependency resolution.
In Nix and Guix, you would indeed not specify that X needs package A version 2.4.2. You could see it as X using the 'build recipe' of A as its dependency. So, if you bump the version of A to 2.4.3, then all packages that use A as a dependency will be rebuilt (or substituted from the a binary cache if the Guix/NixOS build infrastructure has already built the updated packages).
These issues (allow updates within limits) are what I understand as the point of dependency resolution, I'm trying to understand how you do without it.
In such a case e.g. nixpkgs makes two attributes: a_1 and a_2 (and alias a to a_2). Packages that still require A 1.x will have a_1 as one of their dependencies, the rest a.
This is avoided as much as possible, but is sometimes necessary. Common examples are Gtk 2 applications or C/C++ applications that can only be built against a Python 2 interpreter.