코난쌤 블로그
Search
검색
다크 모드
라이트 모드
탐색기
홈
전체 글
카테고리
소개
연락처
개인정보처리방침
태그: inference
5건의 항목
2026년 9월 07일
컨슈머 GPU에서 백만 토큰 에이전트 워크스페이스 돌리기 — KVMem 정리
agent
kv-cache
inference
systems
long-context
2026년 9월 05일
Second Thought: LLM 에이전트의 행동-관찰 구간에서의 병렬 추론 (arXiv 2608.13667)
agent
LLM
inference
parallel-reasoning
latency
2026년 9월 05일
Speculative Macro Commit: 도구 에이전트 지연시간을 19% 줄이는 멀티스텝 투기 실행 (arXiv 2609.03236)
agent
inference
latency
speculative-execution
tool-use
2026년 8월 27일
GLM-5.3-Flash: 320B 모델을 18B 활성 파라미터로 쓰는 법
LLM
GLM
agent
multimodal
coding-agent
inference
long-context
open-weight
2026년 7월 21일
SOPHIA: LLM 추론 루프가 늪에 빠졌을 때 — 숨겨진 활성화 벡터로 탈출시키는 방법
LLM
reasoning
activation-steering
self-loop
inference
agent
interpretability