Docs: How to accelerate question answering (#10179)

### What problem does this PR solve? ### Type of change - [x] Documentation Update
2026-02-02 00:25:06 +08:00 · 2025-09-19 18:18:46 +08:00
parent 6c24ad7966
commit 3f1741c8c6
7 changed files with 89 additions and 23 deletions
--- a/docs/guides/chat/best_practices/accelerate_question_answering.mdx
+++ b/docs/guides/chat/best_practices/accelerate_question_answering.mdx
@ -6,21 +6,22 @@ slug: /accelerate_question_answering
 # Accelerate answering
 import APITable from '@site/src/components/APITable';

-A checklist to speed up question answering.
+A checklist to speed up question answering for your chat assistant.

 ---

 Please note that some of your settings may consume a significant amount of time. If you often find that your question answering is time-consuming, here is a checklist to consider:

- In the **Prompt engine** tab of your **Chat Configuration** dialogue, disabling **Multi-turn optimization** will reduce the time required to get an answer from the LLM.
- In the **Prompt engine** tab of your **Chat Configuration** dialogue, leaving the **Rerank model** field empty will significantly decrease retrieval time.
+- Disabling **Multi-turn optimization** will reduce the time required to get an answer from the LLM.
+- Leaving the **Rerank model** field empty will significantly decrease retrieval time.
+- Disabling the **Reasoning** toggle will reduce the LLM's thinking time. For a model like Qwen3, you also need to add `/no_think` to the system prompt to disable reasoning.
 - When using a rerank model, ensure you have a GPU for acceleration; otherwise, the reranking process will be *prohibitively* slow.

 :::tip NOTE 
 Please note that rerank models are essential in certain scenarios. There is always a trade-off between speed and performance; you must weigh the pros against cons for your specific case.
 :::

- In the **Assistant settings** tab of your **Chat Configuration** dialogue, disabling **Keyword analysis** will reduce the time to receive an answer from the LLM.
+- Disabling **Keyword analysis** will reduce the time to receive an answer from the LLM.
 - When chatting with your chat assistant, click the light bulb icon above the *current* dialogue and scroll down the popup window to view the time taken for each task:  
   ![enlighten](https://github.com/user-attachments/assets/fedfa2ee-21a7-451b-be66-20125619923c)