-
- Replace langchain4j-ollama with langchain4j-open-ai; AgentFactory/ QueryTool depend on ChatModel/StreamingChatModel interfaces - OllamaJsonClient -> LlmJsonClient: raw /v1/chat/completions with response_format json_schema; single OkHttpClient instance - Thinking off at all 3 call sites via reasoning_effort=none (the only switch that works on /v1; think:false and /no_think are ignored) - Drop client-side length params (num_ctx/num_predict/max_tokens) — server-side OLLAMA_CONTEXT_LENGTH=16384 on xlyllm covers them - Unified tracing: TracingChatModelListener.record() shared by listener callbacks and LlmJsonClient; per-request model name from ctx; sql model now has the listener too (was untraced and mislabeled) - Config keys langchain4j.ollama.* -> llm.{base-url,api-key,chat-model, sql-model}; both models = qwen3.6-27b-iq3:latest (tools+thinking+ vision, verified: streaming tool-calls / json_schema / reasoning off) Verified end-to-end on :8199: intent gate traced (340/28 tokens), agent tool loop, correct answers; warm 1-4s, cold load ~96s (KEEP_ALIVE=30s). -
P4 retirement + code slim-down (~18.5k lines deleted), on top of the intent-gate WIP (docs §17/§18 + agent/tool refinements): - Delete old 8-scene multi-agent stack: XlyErpService, SceneSelector/ Chati/ErpAi/DynamicTableNl2Sql agents, DynamicToolProvider (62 meta tools), Scene/ToolMeta/ParamRule entities+mappers+startup caches - Delete whole milvus/tts/ocr packages, /api/tts + /api/ocr endpoints, tts.html, python stream_server.py; strip dead TTS JS from chat.html (playByIndex/handleNormalResponse had no callers) - Three-pass orphan sweep: 17 dead utils, 16 dead entities, dead constants/exceptions, RedisService+RedisConfig, duplicate JacksonConfig, OperableChatMemoryProvider, old prompt generators; fold PageController into MvcConfig; trim OkHttpUtil 495→39 lines - pom: ~35→13 deps (drop mybatis/JPA/webflux/tika/pdfbox/poi/jieba/ jsoup/gson/fastjson2/springdoc/mapstruct/pagehelper/pinyin4j/ spring-retry/jnr-ffi/hutool/ocr/milvus/embeddings; add starter-jdbc); war 200M+→48M - application.yml: drop milvus/tts/ocr/tesseract/mybatis blocks; GlobalExceptionHandler returns plain {code,message} JSON - docs: mark P4 retirement done in agent-architecture §13/§18 Verified: mvn clean package OK; boot smoke on :8199 (Started 0.9s, /chat 200, health db UP). Old /api/tts & /api/ocr now 404 by design.
-
Add a deterministic §5 intent stage before ReAct: extract {intent, form, entities, missing slots} via constrained JSON, then expose only the 3-5 tools relevant to that intent instead of all 12. New: - agent/Intent, agent/ToolScope: intent + visible-tool-set model - service/IntentService: constrained-JSON intent/entity extraction - service/OllamaJsonClient: JSON-mode Ollama calls - service/SlotFillService: slot extraction/prefill of known values Wire through AgentFactory, SystemPromptService, QueryTool, ProposeWriteTool, ErpClient, AgentChatController, OpController, chat.html.
-
…at UI (question/form_collect/token) - QueryTool: enforce table allowlist (viw_* + form data-sources + field-dict tables) via jsqlparser TablesNamesFinder -> blocks NL2SQL reading gdslogininfo/sysjurisdiction/ai_op_queue etc; single-table tenant predicate injection (sBrandsId); brand hint in prompt - TracingChatModelListener: config-gated Langfuse ingestion export (generation span) via JDK HttpClient, no new deps; docker-compose.langfuse.yml self-host - chat.html: send identity+token in chat req; render question(options) + form_collect(rich fields); forward Authorization on confirm/cancel - application-saaslocal.yml: langfuse config (off by default)
-
…ormCollect + Skills + KgSearch - AgentIdentity + AgentFactory: build agent per request with token+form-allowlist carried in tool instances (per-call context; robust vs LC4j callback threading) - ErpClient: token-aware read/write/examine; user-token expiry does NOT silently re-login as dev-admin (no privilege escalation) - AuthzService: devIdentity()/userIdentity() resolve granted-form set (sAuthsId) - proposeExamine tool + OpController examine dispatch + forward user Authorization on confirm - InteractionTool.askUser (structured options), FormCollectTool.collectForm (real gdsconfigformslave schema) - SkillService + SkillTool.loadSkill + ai_skill digest in system prompt - KgQueryTool.kgSearch: L2 neighbors/flow + L3 field->table - tools ErpReadTool/ProposeWriteTool/QueryTool/FormCollectTool now per-request (not @Component)
-
…i_op_queue staging cols + ai_skill registry - proposeCreate tool -> ai_op_queue draft(payload) -> confirm executes addBusinessData - ai_op_queue: add sPayload/sSourceRef/bAutoExecute/sResultBillId/sErrorMsg/tExecutedDate/tExpireAt (create/examine/auto-flow ready) - ai_skill registry table + seed playbooks (Skills §6) - QueryTool: self-repair retry (feed SQL error back to model, up to 3x) - TracingChatModelListener: per-call LLM latency/token/error trace (Langfuse hook point)
-
- FormResolverService: shared KG resolution (entity->master table excl. viw_*, field 中文->column, name field, labels) - ErpReadTool.lookupRecord(entity, record): reliably resolves the master table + returns a single record's full labelled fields — fixes the 'specific field of a specific record' case (e.g. 必胜客的销售员 -> 马艺祖) - AuditService + ai_audit_log: immutable audit of propose/confirm/cancel/query (who/when/action/target/result) - QueryTool.queryData(question): read-only SQL escape hatch — coder-model NL2SQL grounded on 字段字典, jsqlparser single-SELECT guard, blocks OUTFILE/LOAD_FILE/information_schema/SLEEP, forced LIMIT, audited. SAFE by construction; NL2SQL table/join accuracy for complex questions is limited (architecture-flagged, deferred). Verified: lookupRecord works; audit rows written; Query guards + LIMIT + audit work (NL2SQL grounding still mis-picks entity master on multi-table questions).