-
- Replace langchain4j-ollama with langchain4j-open-ai; AgentFactory/ QueryTool depend on ChatModel/StreamingChatModel interfaces - OllamaJsonClient -> LlmJsonClient: raw /v1/chat/completions with response_format json_schema; single OkHttpClient instance - Thinking off at all 3 call sites via reasoning_effort=none (the only switch that works on /v1; think:false and /no_think are ignored) - Drop client-side length params (num_ctx/num_predict/max_tokens) — server-side OLLAMA_CONTEXT_LENGTH=16384 on xlyllm covers them - Unified tracing: TracingChatModelListener.record() shared by listener callbacks and LlmJsonClient; per-request model name from ctx; sql model now has the listener too (was untraced and mislabeled) - Config keys langchain4j.ollama.* -> llm.{base-url,api-key,chat-model, sql-model}; both models = qwen3.6-27b-iq3:latest (tools+thinking+ vision, verified: streaming tool-calls / json_schema / reasoning off) Verified end-to-end on :8199: intent gate traced (340/28 tokens), agent tool loop, correct answers; warm 1-4s, cold load ~96s (KEEP_ALIVE=30s).