Curriculum Vitae (CV)

학력

  • 성균관대학교 빅데이터학과 석사 (2026.03 – 2028.02 졸업 예정)
  • 가천대학교 AI·소프트웨어학부 소프트웨어전공 학사 (2019.03 – 2025.02)
    • 학점 3.81 / 4.5

경력

이노디스AX — 개발팀 주임연구원

  • 2025.03 – 현재 · LLM / Agent / 생성형 AI · Full-stack
  • 담당업무: Agentic AI 기반 검색·챗봇 시스템 연구개발 및 AI 모델 서빙·인프라 구성

  • SIGMOID AI — 생성형 LLM과 Semantic Search를 결합한 RAG 기반 AI 검색·챗봇 시스템 설계·개발
    • OpenSearch + Vector DB 인덱싱 및 RAG 파이프라인 구축
    • 검색 엔진 마이그레이션: Apache 2.0 라이선스 마지막 버전인 Elasticsearch 7.10.2에서 최신 OpenSearch 3.x로 전환 — 인덱스·매핑 이관, 쿼리·클라이언트 호환성 대응
    • 범용 파일 인덱싱 API 개발 (PDF, HWP, PPT 등)
    • LangChain / LangGraph 기반 AI 에이전트 설계 — 고정된 워크플로(predefined path)가 아닌, 모델이 자율 판단으로 툴 호출을 이어가는 동적(agentic) 오케스트레이션 구현
    • Multi-Agent + 페르소나 기반 역할 분리로 응답 일관성·역할 유지
    • 배포 후 지속 모니터링으로 검색 품질 측정, 시간 경과에 따른 성능 저하(drift) 탐지·대응
    • 성과: 고객사 DB 연결 없이 범용 파일 인덱싱 실현, 다국어 입출력 지원, 최근 3개월 평균 3.8k 사용자
  • SIGMOID LLM Manager — 도메인·상태 기반 AI 모델 자동 라우팅 LLM Gateway
    • FastAPI 기반 API 인증 서버 및 모델 자동 라우팅 구현 — 모든 LLM 트래픽이 거쳐 가는 단일 게이트웨이(API 키 인증, IP 접근 제어, OpenAI 호환 프록시)
    • Prometheus + Grafana 기반 모니터링·시각화
    • 성과: 모델 상태(응답 지연·오류율) 기반 Failover 및 우선순위 라우팅, 동시 요청 100건 0.1초 내 처리, 월 평균 3만 건 API
  • 공통 인프라 · 운영 — 사내 서버망과 클라우드 망을 함께 관리하며 서비스 인프라 구축·운영
    • 네트워크: 서버 간 네트워크 구성·연결, 방화벽·ACL 정책 관리, 도메인/DNS 관리, SSL/TLS 인증서 발급·갱신
    • 배포: Docker / docker-compose 기반 다중 서비스 구성·배포, Nginx 웹서버·로드밸런싱 구성으로 무중단 배포, GitLab CI · Jenkins 기반 CI/CD 파이프라인 구축, GPU 서버 세팅(드라이버·CUDA·NVIDIA Container Toolkit)
    • 운영·관측: 로그 수집·관리, 장애 알림 룰 구성, 리소스 모니터링 및 용량 산정, DB 백업·복구
    • 클라우드: 네이버 클라우드(NCP) 리소스(VPC·서버·스토리지) 관리 및 온프렘↔클라우드 연동, CSAP·GS 인증 취득 지원, KMS v2 키 관리
    • AI 서비스: vLLM 기반 로컬 LLM·Reranker 추론 서빙 인프라 구축·운영 — H100 ×1 1대 + RTX 4090 ×2 3대(총 4대) GPU 서버에 모델 서빙 구성, 외부 API 없이 자체 서빙
    • AI 가속기 도입 벤치마크 (단독 담당): Pliops SSD 레벨 KV 캐시 오프로딩 가속기에 대해 국내 유통사의 수입 검토를 위한 성능 검증을 전담 — vLLM 추론 서빙 환경에서 TTFT·프리필 지연, 처리량, KV 캐시 히트율·용량을 측정하는 벤치마크 시나리오 설계·수행 및 결과 리포팅
    • 폐쇄망 로컬 LLM 서빙 검증: 폐쇄망 환경 도입을 위한 로컬 LLM 서빙 구축 — RTX 4090 ×1 / ×2, MacBook Pro 48GB, DGX Spark 하드웨어별 벤치마크 및 실사용 테스트로 도입 구성 비교·검증, Linux · CUDA 기반 AI 인프라 세팅 포함

슈어소프트테크 — AI기술팀 인턴

  • 2024.07 – 2024.11 · AI / NLP / sLLM / Agent
  • 담당업무: 도메인 특화 sLLM 파인튜닝 · LLMOps · 서빙 및 LLM 평가 벤치마크 연구개발

  • Code2Chart — 도메인 특화 sLLM 파인튜닝
    • 데이터 정제·수집, 학습 파이프라인 구축, Docker 기반 sgLang 서빙
    • LLMOps: MLflow · WandB · DVC로 튜닝 데이터셋 · 하이퍼파라미터 레시피 · 모델 아티팩트를 버전 관리 — 실험 이력 추적 및 재현 가능한 학습 파이프라인 구성
    • SFT 기법(FFT · LoRA · QLoRA) 비교 및 양자화(Quantization) 기반 경량 파인튜닝, CoT 데이터 학습
    • H100 80GB × 8 단일 노드에서 FSDP · DeepSpeed ZeRO 샤딩으로 분산 학습 환경 구성 — LLaMA-Factory 기반 학습 구성에 Mixed Precision · Gradient Checkpointing · FlashAttention · FlashInfer 커널 백엔드 적용
    • vLLM Prefill/Decode 분리(P/D Disaggregation) 서빙 구성 — KV 캐시 전송 경로 및 인스턴스별 병렬도·배치 정책 분리 조정
    • 성과: Base 모델 33점 → 파인튜닝 67점 (2배), 사내 배포 및 QA 실사용 검증
  • sLLM Benchmark — 도메인 특화 LLM-as-a-Judge 벤치마크
    • Chain-of-Thought(CoT) 기반 평가 템플릿, 10개 세부 항목 설계
    • 성과: GT 데이터 100개 기준 신뢰도 89.5%
  • 기타: LangChain 기반 태스크 자동화 에이전트 프로토타입 개발

AI 활용 역량

  • Claude Code 기반 업무 자동화 — AI 코딩 에이전트를 실무 전반에 도입해 반복 작업을 자동화하고 개발 생산성을 대폭 향상 (Claude Code 메인 + Codex 서브 멀티모델을 Orca ADE(Agentic Development Environment) 환경에서 활용)
    • 커스텀 스킬 개발·공유: 사내에서 자주 쓰는 업무를 스킬(Skill)로 직접 제작해 팀에 공유·활용
    • 에이전트 하네스 구축: oh-my-claudecode 등 플러그인·스킬·서브에이전트를 조합해 멀티에이전트 오케스트레이션 환경 구성 — 구현 → QA · 테스트 · 검증까지 자동으로 수행되는 파이프라인 운영
    • Hooks 기반 정책 자동화: 훅(Hooks)으로 개발 정책·점검 규칙을 자동 강제
    • 토큰 최적화: 토큰 최적화 프록시(RTK) 를 실무에 적용해 토큰 사용량·비용 절감
    • 컨텍스트 엔지니어링: CLAUDE.md로 프로젝트 공용 지침·개발 방식·점검 방식을 정의해 에이전트의 정확도와 일관성 확보
    • AI 활용 Full-stack 확장: 백엔드·AI 중심 업무를 AI 코딩 에이전트를 활용해 프론트엔드까지 Full-stack으로 확장, 서비스 개발 전 영역 겸임
  • 사내 AI 확산 주도 — 기존 개발 지식(인프라·CI/CD·테스트)과 AI 에이전트를 결합한 노하우를 조직에 전파
    • 이노디스AX에서 Claude Code 활용 사내 세미나 개최 — 토큰 최적화 프록시, codebase-mcp 등 실전 활용법 공유
    • 프로젝트별 표준 가이드·설정 구축: 프로젝트마다 CLAUDE.md·공용 설정·개발 표준 가이드를 정비해 팀 단위 AI 도입 기반 마련
    • 비개발 직군 대상 AX 교육 — 현업 부서가 업무에 AI를 활용할 수 있도록 교육·지원

프로젝트

프로젝트 분야
GuardianWatch 비전 AI 기반 아동 안전 모니터링 (졸업작품). CCTV 실시간 다중 객체 추적(MOT)과 행동 분석으로 위험 상황을 감지하고, 아이의 위치·행동·위험 알림을 부모 앱에 실시간 전송
resumable-upload 의존성 없이 Python으로 구현한 TUS 재개 가능 업로드 프로토콜 (v1.0.0) — 서버 & 클라이언트
Distributed HAR System for Scalable Wi-Fi Sensing 프라이버시 보존 인간 행동 인식 (INC Lab, 2023.03–2024.08). 라즈베리파이 4B + Nexmon CSI 기반 분산 수집 노드, Wi-Fi CSI 신호를 스펙트로그램 이미지로 변환ViT로 분류 — 데이터의 1/N만으로 정확도 90%, 도청 모델 정확도는 ~30% 이하로 억제

오픈소스

기여 내역

프로젝트 기여 내용
vllm-project/vllm GGUF 모델 로딩 경로 개선 및 오프라인·클라우드 스토리지 로딩 버그 수정 — 고성능 LLM 추론 엔진 #39471 · #39155 · #29582 · #29137
vllm-project/vllm-metal Metal paged 경로의 구조화 출력(grammar / structured output) 지원 구현 — Apple Silicon용 vLLM 하드웨어 플러그인 #280
hiyouga/LLaMA-Factory 학습 설정 로딩을 OmegaConf 기반으로 개선 — 통합 LLM 파인튜닝 프레임워크 #8505
tus/tus.io 직접 구현한 resumable-upload를 공식 구현체 목록에 등재 — TUS 재개 가능 업로드 프로토콜 공식 사이트 #603 · #545
vibrantlabsai/ragas 설치 및 퀵스타트 문서 오류 수정 — LLM 애플리케이션 평가 프레임워크 #2564 · #2563
langflow-ai/langflow 한국어 README 추가 — 비주얼 LLM 워크플로우 빌더 #2883
OpenBMB/ChatDev GPT-4o / GPT-4o-mini 모델 지원 추가, 한국어 README 보완 — 멀티에이전트 소프트웨어 개발 #405 · #404
OpenBMB/RepoAgent RAG 모듈의 시스템 프롬프트 오류 수정 — LLM 기반 저장소 문서화 #74
stitionai/devika GPT-4o-mini 모델 지원 추가 — 에이전트형 AI 소프트웨어 엔지니어 #630
context-labs/autodoc GPT-4o / GPT-4o-mini 모델 지원 추가 — LLM 기반 코드베이스 문서화 #46
pqrs-org/KE-complex_modifications 기존 단축키를 유지하는 원화(₩) 입력 키 매핑 규칙 추가 — Karabiner-Elements 개발자 도구 #1809

연구 및 활동

  • Pseudo Lab
  • INC Lab (무선통신), 가천대학교 (2023.03 – 2024.08)
    • WiFi CSI 기반 인간 행동 인식, 프라이버시 보존 분산 센싱
    • 라즈베리파이 4B + Nexmon CSI로 CSI 수집 노드를 구성해 분산 센싱 환경 구축 및 데이터 수집
    • 논문: Distributed HAR System for Scalable Wi-Fi Sensing — 확장 가능한 WiFi 센싱을 위한 분산 행동 인식 시스템 (CSI 스펙트로그램 → ViT 분류, 데이터 1/N로 정확도 90%, 도청 모델 정확도 ~30% 이하)
  • BDA (Big Data Association, 現 BDAI) (2024.03 – 2024.08)
    • EDA, 통계 분석, NLP 전처리

자격증 및 수상

  • 정보처리기사 (2024.09)
  • 데이터분석준전문가 (ADsP) (2024.06)
  • SQL개발자 (SQLD) (2024.04)
  • TOEIC Speaking 150 — Intermediate High (2024.08)
  • 🥈 K-디지털 챌린지: NET Challenge Camp S10 은상 (과학기술정보통신부, 2023.12)

Education

  • M.S. in Department of Big Data, Sungkyunkwan University (2026.03 – 2028.02, expected)
  • B.S. in Department of Software, School of AI·Software, Gachon University (2019.03 – 2025.02)
    • GPA 3.81 / 4.5

Experience

innodisAX — Associate Research Engineer, Development Team

  • 2025.03 – Present · LLM / Agent / Generative AI · Full-stack
  • Scope: R&D on agentic-AI search and chatbot systems, and AI model serving & infrastructure

  • SIGMOID AI — Designed and built a RAG-based AI search and chatbot system combining generative LLMs with semantic search
    • Built OpenSearch + Vector DB indexing and a RAG pipeline
    • Search engine migration: moved from Elasticsearch 7.10.2, the last Apache-2.0-licensed release, to the latest OpenSearch 3.x — index/mapping migration and query/client compatibility work
    • Developed a general-purpose file indexing API (PDF, HWP, PPT, etc.)
    • Designed AI agents with LangChain / LangGraphdynamic, agentic orchestration where the model autonomously directs its own tool calls, rather than a predefined workflow path
    • Role separation via multi-agent + persona design for consistent responses and stable role adherence
    • Continuous post-deployment monitoring to measure search quality and counter performance drift over time
    • Impact: general-purpose file indexing without connecting to customer DBs, multilingual I/O, avg. 3.8k users over the last 3 months
  • SIGMOID LLM Manager — an LLM Gateway for domain- and status-based automatic AI model routing
    • Implemented a FastAPI-based API auth server and automatic model routing — a single gateway all LLM traffic passes through (API key auth, IP access control, OpenAI-compatible proxy)
    • Monitoring and visualization with Prometheus + Grafana
    • Impact: failover and priority routing based on model status (latency, error rate), 100 concurrent requests handled within 0.1s, ~30k API calls/month
  • Shared infrastructure & operations — managed both the on-premise server network and the cloud network, building and running the service infrastructure
    • Networking: server-to-server network setup, firewall / ACL policy management, domain & DNS management, SSL/TLS certificate issuance and renewal
    • Deployment: multi-service setup and deployment with Docker / docker-compose, zero-downtime deployment via Nginx web servers and load balancing, CI/CD pipelines with GitLab CI / Jenkins, GPU server provisioning (drivers, CUDA, NVIDIA Container Toolkit)
    • Operations & observability: log collection and retention, alerting rules for incidents, resource monitoring and capacity planning, DB backup and recovery
    • Cloud: Naver Cloud (NCP) resource management (VPC, servers, storage) and on-prem ↔ cloud integration, supported CSAP / GS certification, KMS v2 key management
    • AI services: built and operated the vLLM-based local LLM / Reranker inference serving infrastructure — model serving across four GPU servers — one with 1× H100 and three with 2× RTX 4090, served in-house without external APIs
    • AI accelerator benchmarking (sole owner): led the performance validation for a Korean distributor’s import evaluation of Pliops SSD-level KV-cache offload accelerators — designed and ran the benchmark scenarios on vLLM inference serving (TTFT / prefill latency, throughput, KV-cache hit rate and capacity) and reported the results
    • Air-gapped local LLM serving validation: built local LLM serving for closed-network deployment — benchmarked and real-use-tested candidate setups across RTX 4090 ×1 / ×2, MacBook Pro 48GB, and DGX Spark hardware, including Linux / CUDA-based AI infrastructure provisioning

Suresofttech — AI Technology Team Intern

  • 2024.07 – 2024.11 · AI / NLP / sLLM / Agent
  • Scope: Domain-specific sLLM fine-tuning, LLMOps, serving, and LLM evaluation benchmark R&D

  • Code2Chart — domain-specific sLLM fine-tuning
    • Data cleaning/collection, training pipeline, Docker-based sgLang serving
    • LLMOps: versioned tuning datasets, hyperparameter recipes, and model artifacts with MLflow · WandB · DVC — experiment tracking and a reproducible training pipeline
    • Compared SFT methods (FFT · LoRA · QLoRA) and applied quantization for lightweight fine-tuning, trained on Chain-of-Thought (CoT) data
    • Distributed training on a single node with 8× H100 80GBFSDP / DeepSpeed ZeRO sharding, with mixed precision, gradient checkpointing, and FlashAttention / FlashInfer kernel backends applied in a LLaMA-Factory training setup
    • vLLM prefill/decode disaggregated (P/D) serving — KV-cache transfer path and per-instance parallelism/batching policies tuned separately
    • Impact: base model 33 → fine-tuned 67 (2×), internal deployment and QA real-use validation
  • sLLM Benchmark — domain-specific LLM-as-a-Judge benchmark
    • Chain-of-Thought (CoT) evaluation templates, 10 detailed criteria
    • Impact: 89.5% reliability on 100 ground-truth samples
  • Other: prototyped a LangChain-based task automation agent

AI-Assisted Development

  • Claude Code-driven workflow automation — adopted AI coding agents across day-to-day work, automating repetitive tasks and significantly boosting development productivity (Claude Code as the primary agent and Codex as a secondary model, inside the Orca ADE — Agentic Development Environment)
    • Custom skill development & sharing: built frequently-used internal workflows into reusable Skills and shared them across the team
    • Agent harness engineering: composed plugins, skills, and subagents (e.g. oh-my-claudecode) into a multi-agent orchestration setup — a pipeline that runs implementation through QA, testing, and verification automatically
    • Hook-based policy automation: enforced development and review policies automatically via hooks
    • Token optimization: applied a token-optimizing proxy (RTK) in production use, cutting token usage and cost
    • Context engineering: defined project-wide conventions, development practices, and review criteria in CLAUDE.md for accurate, consistent agent behavior
    • AI-augmented full-stack: extended a backend/AI-focused role into full-stack development by leveraging AI coding agents, covering the frontend side of services as well
  • Driving internal AI adoption — spread know-how combining existing engineering skills (infrastructure, CI/CD, testing) with AI agents across the organization
    • Hosted an in-house Claude Code seminar at innodisAX — covering the token-optimizing proxy, codebase-mcp, and hands-on usage patterns
    • Per-project standard guides & configuration: set up CLAUDE.md, shared settings, and development standards for each project, laying the groundwork for team-wide AI adoption
    • AX training for non-developer roles — helped business teams apply AI to their daily work

Projects

Project Domain
GuardianWatch vision-AI childcare safety monitoring (graduation project). Real-time multi-object tracking (MOT) and behavior analysis on CCTV to detect hazards and push children’s location, behavior, and danger alerts to a parent app
resumable-upload Zero-dependency Python implementation of the TUS resumable upload protocol (v1.0.0) — server & client.
Distributed HAR System for Scalable Wi-Fi Sensing privacy-preserving human activity recognition (INC Lab, 2023.03–2024.08). Distributed capture nodes on Raspberry Pi 4B + Nexmon CSI; Wi-Fi CSI signals turned into spectrogram images and classified with a ViT90% accuracy using only 1/N of the data; eavesdropper model accuracy kept below ~30%

Open Source

Contributions

Project Contribution
vllm-project/vllm Improved the GGUF model loading path; fixed offline and cloud-storage loading bugs — High-throughput LLM inference engine #39471 · #39155 · #29582 · #29137
vllm-project/vllm-metal Implemented grammar / structured-output support for the Metal paged path — Hardware plugin for vLLM on Apple Silicon #280
hiyouga/LLaMA-Factory Switched training-config loading to OmegaConf — Unified LLM fine-tuning framework #8505
tus/tus.io Listed my own resumable-upload on the official implementations page — Official site for the TUS resumable upload protocol #603 · #545
vibrantlabsai/ragas Fixed errors in the install and quickstart documentation — Evaluation framework for LLM applications #2564 · #2563
langflow-ai/langflow Added a Korean README — Visual LLM workflow builder #2883
OpenBMB/ChatDev Added GPT-4o / GPT-4o-mini model support and improved the Korean README — Multi-agent software development #405 · #404
OpenBMB/RepoAgent Fixed a broken system prompt in the RAG module — LLM-powered repository documentation #74
stitionai/devika Added GPT-4o-mini model support — Agentic AI software engineer #630
context-labs/autodoc Added GPT-4o / GPT-4o-mini model support — LLM-based codebase documentation #46
pqrs-org/KE-complex_modifications Added a Korean-won key mapping rule that preserves the existing shortcut — Karabiner-Elements developer tooling #1809

Research & Activities

  • Pseudo Lab
  • INC Lab (Wireless Communication), Gachon University (2023.03 – 2024.08)
    • WiFi CSI-based human activity recognition, privacy-preserving distributed sensing
    • Built the distributed sensing testbed and collected data with Raspberry Pi 4B + Nexmon CSI capture nodes
    • Paper: Distributed HAR System for Scalable Wi-Fi Sensing — distributed activity recognition for scalable Wi-Fi sensing (CSI spectrograms classified by a ViT, 90% accuracy with 1/N of the data, eavesdropper accuracy kept below ~30%)
  • BDA (Big Data Association, now BDAI) (2024.03 – 2024.08)
    • EDA, statistical analysis, NLP preprocessing

Certifications & Awards

  • Engineer Information Processing (2024.09)
  • Advanced Data Analytics Semi-Professional (ADsP) (2024.06)
  • SQL Developer (SQLD) (2024.04)
  • TOEIC Speaking 150 — Intermediate High (2024.08)
  • 🥈 K-Digital Challenge: NET Challenge Camp S10 — Silver Prize (Ministry of Science and ICT, 2023.12)

Skills

구분Category 스킬Skills
LLM / NLP LangChainLangGraphRAGRerankerLLM-as-a-JudgeAgent
Fine-tuning SFTFFTLoRAQLoRAQuantizationLLaMA-Factory
LLM Serving vLLMsgLangTGI
Backend PythonFastAPIDjangoFlask
ML / DL PyTorchTransformersCUDAComputer VisionOpenCV
MLOps DVCWandBMLflow
Data / Search ElasticsearchOpenSearchMySQLSQLPandasMATLAB
Infra LinuxDockerGitNginxGitLab CIJenkinsGitHub ActionsNaver Cloud (NCP)PrometheusGrafana