Limitations with output tokens that prevent detailed answers #952

orangeelephant2013 · 2024-08-09T07:24:09Z

Hello, I had raised this in the discussions but received no answers (link: #936 (comment)). How can we ensure that answers aren't truncated baed not he maximum output tokens? Even if the tokens are set at 2000, the answers are truncated when using the bling-phi-3-gguf model (and I assume any other model dependent on llmware's apis). I tried multiple methods to change the maximum output tokens including passing it into the method to get the model as well as in the gguf config. Can someone please help. thanks!

NYDocutest · 2024-08-12T15:06:12Z

Hi @orangeelephant2013 Thank you for bringing this to our attention. We made the necessary update in our repo as shown so please try again with a fresh clone.

orangeelephant2013 · 2024-08-13T06:39:43Z

Thanks so much for the reply. We will update and try again. I will keep you posted once things work as expected. thanks again

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Limitations with output tokens that prevent detailed answers #952

Limitations with output tokens that prevent detailed answers #952

orangeelephant2013 commented Aug 9, 2024

NYDocutest commented Aug 12, 2024

orangeelephant2013 commented Aug 13, 2024

Limitations with output tokens that prevent detailed answers #952

Limitations with output tokens that prevent detailed answers #952

Comments

orangeelephant2013 commented Aug 9, 2024

NYDocutest commented Aug 12, 2024

orangeelephant2013 commented Aug 13, 2024