Eric Hare 194f42be81 perf: cache Ollama model capabilities to fix slow Cloud model toggles (#13722)
* perf: cache Ollama model capabilities to fix slow Cloud model toggles

Enabling/disabling models for an Ollama provider pointed at the cloud
base URL (https://ollama.com) lagged on every toggle, while a local
Ollama stayed instant (issue #12399).

Root cause: get_ollama_models() probes capabilities with one POST
/api/show per model on top of the GET /api/tags listing, so a single
catalog read costs N + 1 upstream requests. GET /enabled_models reads
both llm and embeddings (replace_with_live_models(model_type=None)), so
each read paid 2 x (N + 1), and every model toggle triggers a refetch.
On Ollama Cloud's large public catalog over the internet this crawls;
local Ollama has a tiny catalog and ~0 latency so it stays fast. The
existing 30s list cache only helps clustered same-capability reads.

Fix: add a per-model capability cache keyed by (base_url, model_name).
A model:tag's capabilities are intrinsic to the model, so the /api/show
result is reused instead of re-probed. This makes the llm and embedding
reads share a single fan-out (N probes, not 2N) and makes any read after
the short list-TTL re-probe only models never seen before, not the whole
catalog. TTL is bounded at 10 minutes so a re-pull that changes a tag's
capability class self-heals quickly. A probe failure is still absorbed
and left uncached so one bad model never poisons or sticks in the
catalog.

Backend-only and behavior-preserving: the model lists returned are
identical; only the upstream request count drops.

* [autofix.ci] apply automated fixes

* [autofix.ci] apply automated fixes (attempt 2/3)

* perf: avoid caching empty Ollama capabilities and prune stale entries

Address review feedback on the per-model capability cache (#13722).

A 200 /api/show with no capabilities resolved to [] and was cached for the full TTL, while a real RequestError/HTTPStatusError was correctly left uncached, so a transient empty response could hide a model from the picker for 10 minutes. Cache only a populated capability list, so an empty/absent array is re-probed on the next read like the error path. Real Ollama always returns a populated array, so no legitimately-capable model is ever re-probed.

The capability cache expired on read but never evicted, retaining one entry per distinct model seen for the process lifetime. Prune it to the live catalog on each fresh read so it tracks the current catalog; entries for other base URLs are untouched.

Adds behavioral tests for both paths.

---------

Co-authored-by: autofix-ci[bot] <114827586+autofix-ci[bot]@users.noreply.github.com>
2026-06-19 19:37:17 +00:00
2026-06-09 13:16:48 -07:00
2026-06-09 13:16:48 -07:00
2025-03-20 00:05:55 +00:00
2026-04-23 17:49:53 -07:00
2024-06-04 09:26:13 -03:00
2026-06-19 18:43:48 +00:00

Langflow logo

Release Notes PyPI - License PyPI - Downloads Twitter YouTube Channel Discord Server Ask DeepWiki

Langflow is a powerful platform for building and deploying AI-powered agents and workflows. It provides developers with both a visual authoring experience and built-in API and MCP servers that turn every workflow into a tool that can be integrated into applications built on any framework or stack. Langflow comes with batteries included and supports all major LLMs, vector databases and a growing library of AI tools.

Highlight features

  • Visual builder interface to quickly get started and iterate.
  • Source code access lets you customize any component using Python.
  • Interactive playground to immediately test and refine your flows with step-by-step control.
  • Multi-agent orchestration with conversation management and retrieval.
  • Deploy as an API or export as JSON for Python apps.
  • Deploy as an MCP server and turn your flows into tools for MCP clients.
  • Observability with LangSmith, LangFuse and other integrations.
  • Enterprise-ready security and scalability.

🖥️ Langflow Desktop

Langflow Desktop is the easiest way to get started with Langflow. All dependencies are included, so you don't need to manage Python environments or install packages manually. Available for Windows and macOS.

📥 Download Langflow Desktop

Quickstart

Requires Python 3.103.14 and uv (recommended package manager).

Install

From a fresh directory, run:

uv pip install langflow -U

The latest Langflow package is installed. For more information, see Install and run the Langflow OSS Python package.

Run

To start Langflow, run:

uv run langflow run

Langflow starts at http://127.0.0.1:7860.

That's it! You're ready to build with Langflow! 🎉

📦 Other install options

Run from source

If you've cloned this repository and want to contribute, run this command from the repository root:

make run_cli

For more information, see DEVELOPMENT.md.

Docker

Start a Langflow container with default settings:

docker run -p 7860:7860 langflowai/langflow:latest

Langflow is available at http://localhost:7860/. For configuration options, see the Docker deployment guide.

🛡️ Security

For security information, see our Security Policy.

🚀 Deployment

Langflow is completely open source and you can deploy it to all major deployment clouds. To learn how to deploy Langflow, see our Langflow deployment guides.

Stay up-to-date

Star Langflow on GitHub to be instantly notified of new releases.

Star Langflow

👋 Contribute

We welcome contributions from developers of all levels. If you'd like to contribute, please check our contributing guidelines and help make Langflow more accessible.


Star History Chart

❤️ Contributors

langflow contributors

Description
Langflow is a powerful tool for building and deploying AI-powered agents and workflows.
Readme MIT 2.3 GiB
Languages
Python 64.5%
TypeScript 23.4%
JavaScript 11.4%
CSS 0.3%
Makefile 0.2%
Other 0.1%