You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Hello, I had raised this in the discussions but received no answers (link: #936 (comment)). How can we ensure that answers aren't truncated baed not he maximum output tokens? Even if the tokens are set at 2000, the answers are truncated when using the bling-phi-3-gguf model (and I assume any other model dependent on llmware's apis). I tried multiple methods to change the maximum output tokens including passing it into the method to get the model as well as in the gguf config. Can someone please help. thanks!
The text was updated successfully, but these errors were encountered:
Hi @orangeelephant2013 Thank you for bringing this to our attention. We made the necessary update in our repo as shown so please try again with a fresh clone.
Hello, I had raised this in the discussions but received no answers (link: #936 (comment)). How can we ensure that answers aren't truncated baed not he maximum output tokens? Even if the tokens are set at 2000, the answers are truncated when using the bling-phi-3-gguf model (and I assume any other model dependent on llmware's apis). I tried multiple methods to change the maximum output tokens including passing it into the method to get the model as well as in the gguf config. Can someone please help. thanks!
The text was updated successfully, but these errors were encountered: