• - Replace langchain4j-ollama with langchain4j-open-ai; AgentFactory/
      QueryTool depend on ChatModel/StreamingChatModel interfaces
    - OllamaJsonClient -> LlmJsonClient: raw /v1/chat/completions with
      response_format json_schema; single OkHttpClient instance
    - Thinking off at all 3 call sites via reasoning_effort=none (the only
      switch that works on /v1; think:false and /no_think are ignored)
    - Drop client-side length params (num_ctx/num_predict/max_tokens) —
      server-side OLLAMA_CONTEXT_LENGTH=16384 on xlyllm covers them
    - Unified tracing: TracingChatModelListener.record() shared by listener
      callbacks and LlmJsonClient; per-request model name from ctx; sql
      model now has the listener too (was untraced and mislabeled)
    - Config keys langchain4j.ollama.* -> llm.{base-url,api-key,chat-model,
      sql-model}; both models = qwen3.6-27b-iq3:latest (tools+thinking+
      vision, verified: streaming tool-calls / json_schema / reasoning off)
    
    Verified end-to-end on :8199: intent gate traced (340/28 tokens), agent
    tool loop, correct answers; warm 1-4s, cold load ~96s (KEEP_ALIVE=30s).
    zichun authored
     
    Browse Code »