Skip to content

Navigation Menu

Sign in
Appearance settings

Search code, repositories, users, issues, pull requests...

Provide feedback

We read every piece of feedback, and take your input very seriously.

Saved searches

Use saved searches to filter your results more quickly

Appearance settings
Discussion options

I might have missed something but is it possible to unload a model if it wasn't used for X minutes? Ollama has something like that, freeing up vram for other things (image generation etc.) if the llm isn't in use currently.

You must be logged in to vote

Replies: 1 comment

Comment options

I forgot to add, my question is about the llama-server that you can spin up to create an openapi compatible endpoint.

You must be logged in to vote
0 replies
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Category
🙏
Q&A
Labels
None yet
1 participant
Morty Proxy This is a proxified and sanitized view of the page, visit original site.