ForcePilot/backend/package/yuxi/channels/adapters/wechat/chunker.py
Kris 05abecc02b feat(wechat): 新增完整微信渠道适配器实现
该提交实现了支持企业微信、微信公众号、个人微信桥接三种模式的完整微信渠道适配器,包含以下核心模块:
1. 基础认证与配置相关:auth_adapter、config_reload、setup_contract等
2. 消息处理与格式转换:format、attachment_adapter、outbound_adapter等
3. 多模式客户端支持:wecom/mp子模块,包含加解密、消息收发能力
4. 辅助能力:限速器、防抖、会话绑定、事件映射、模板渲染等
5. 扩展能力:二维码登录、消息读取、特权用户、心跳监控等

实现了完整的微信生态对接能力,支持消息收发、事件处理、API调用限流、配置热重载等功能。
2026-05-12 00:51:04 +08:00

85 lines
2.4 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

from __future__ import annotations
from typing import Literal
NATURAL_BOUNDARIES = [
"\n\n",
"\n",
"",
"",
"",
"",
"",
". ",
"! ",
"? ",
]
class WeChatChunker:
def chunk_text(self, text: str, limit: int = 2048) -> list[str]:
if len(text) <= limit:
return [text]
chunks: list[str] = []
remaining = text
while len(remaining) > limit:
cut = self._find_boundary(remaining, limit)
if cut <= 0:
cut = limit
chunks.append(remaining[:cut])
remaining = remaining[cut:]
if remaining:
chunks.append(remaining)
return chunks
def chunk_markdown(self, text: str, limit: int = 2048) -> list[str]:
if len(text) <= limit:
return [text]
chunks: list[str] = []
remaining = text
code_block = False
while len(remaining) > limit:
if "```" in remaining[:limit]:
fence_idx = remaining.find("```")
if fence_idx >= 0:
code_block = not code_block
if code_block:
close_idx = remaining.find("```", fence_idx + 3)
block_end = close_idx + 3 if close_idx >= 0 else limit
if block_end <= limit:
cut = block_end
else:
cut = limit
else:
cut = self._find_boundary(remaining, limit)
chunks.append(remaining[:cut])
remaining = remaining[cut:]
continue
cut = self._find_boundary(remaining, limit)
if cut <= 0:
cut = limit
chunks.append(remaining[:cut])
remaining = remaining[cut:]
if remaining:
chunks.append(remaining)
return chunks
def chunk(self, text: str, limit: int, mode: Literal["text", "markdown"] = "text") -> list[str]:
if mode == "markdown":
return self.chunk_markdown(text, limit)
return self.chunk_text(text, limit)
@staticmethod
def _find_boundary(text: str, limit: int) -> int:
best = -1
for boundary in NATURAL_BOUNDARIES:
idx = text.rfind(boundary, 0, limit)
if idx > best:
best = idx
return best + 1 if best >= 0 else 0