코난쌤 블로그

홈전체 글카테고리소개연락처개인정보처리방침

태그: inference

5건의 항목

  • 2026년 9월 07일

    컨슈머 GPU에서 백만 토큰 에이전트 워크스페이스 돌리기 — KVMem 정리

    • agent
    • kv-cache
    • inference
    • systems
    • long-context
  • 2026년 9월 05일

    Second Thought: LLM 에이전트의 행동-관찰 구간에서의 병렬 추론 (arXiv 2608.13667)

    • agent
    • LLM
    • inference
    • parallel-reasoning
    • latency
  • 2026년 9월 05일

    Speculative Macro Commit: 도구 에이전트 지연시간을 19% 줄이는 멀티스텝 투기 실행 (arXiv 2609.03236)

    • agent
    • inference
    • latency
    • speculative-execution
    • tool-use
  • 2026년 8월 27일

    GLM-5.3-Flash: 320B 모델을 18B 활성 파라미터로 쓰는 법

    • LLM
    • GLM
    • agent
    • multimodal
    • coding-agent
    • inference
    • long-context
    • open-weight
  • 2026년 7월 21일

    SOPHIA: LLM 추론 루프가 늪에 빠졌을 때 — 숨겨진 활성화 벡터로 탈출시키는 방법

    • LLM
    • reasoning
    • activation-steering
    • self-loop
    • inference
    • agent
    • interpretability

Created with Quartz v4.5.2 © 2026

  • 소개
  • 연락처
  • 개인정보처리방침
  • 전체 글