Files
project_6/upstream_ref/xllm/xllm/parser/reasoning_detector.h
EX Engine 002f9879b2 ref(upstream): FULL TREE — Deep-Spark xllm (1470) + ds_vllm csrc/models (703)
Replaces cherry-picked upstream_ref with complete source trees.

xllm/ — Iluvatar official C++ inference engine (15MB, 1470 files)
  Complete: kernels → layers → models → runtime → scheduler → api
  Excluded: .git, binary images, third_party submodule checkouts

ds_vllm/ — Iluvatar official vllm fork (8MB, 703 files)
  Included: csrc/ (ALL CUDA kernels), fused_moe/, qwen3_5 model, _custom_ops
  Excluded: tests, benchmarks, docs, examples (not needed for reference)

Critical call chains now fully traceable:
  MoE: moe_topk_softmax_kernels.cuh → ixformer.h → fused_moe.cpp → layer
  GDN: qwen3_gated_delta_net_base.cpp → qwen3_5_gated_delta_net.cpp
  Attention: ixformer.h → xllm_paged_attention → attention.cpp
2026-08-10 02:54:03 +00:00

65 lines
2.2 KiB
C++

/* Copyright 2025 The xLLM Authors. All Rights Reserved.
Copyright 2024 The ScaleLLM Authors. All Rights Reserved.
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at
https://github.com/jd-opensource/xllm/blob/main/LICENSE
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.
==============================================================================*/
#pragma once
#include <optional>
#include <string>
namespace xllm {
struct ReasoningResult {
std::optional<std::string> normal_text = std::nullopt;
std::optional<std::string> reasoning_text = std::nullopt;
ReasoningResult() = default;
ReasoningResult(std::optional<std::string> normal,
std::optional<std::string> reasoning)
: normal_text(normal), reasoning_text(reasoning) {}
};
class ReasoningDetector {
public:
ReasoningDetector(const std::string& think_start_token,
const std::string& think_end_token,
bool force_reasoning = false,
bool stream_reasoning = true);
~ReasoningDetector() = default;
// Detects and parses reasoning sections in the provided text. Returns both
// reasoning content and normal text separately.
ReasoningResult detect_and_parse(std::string& text);
// Streaming incremental parsing for reasoning content.
// Handles partial reasoning tags and content.
//
// If stream_reasoning is False:
// Accumulates reasoning content until the end tag is found
// If stream_reasoning is True:
// Streams reasoning content as it arrives
ReasoningResult parse_streaming_increment(std::string& new_text);
protected:
std::string think_start_token_;
std::string think_end_token_;
bool in_reasoning_;
bool stream_reasoning_;
std::string buffer_ = "";
bool stripped_think_start_ = false;
};
} // namespace xllm