Curriculum Vitae (CV)
학력
- 성균관대학교 빅데이터학과 석사 (2026.03 – 2028.02 졸업 예정)
- 가천대학교 AI·소프트웨어학부 소프트웨어전공 학사 (2019.03 – 2025.02)
- 학점 3.81 / 4.5
경력
이노디스AX — 개발팀 주임연구원
- 2025.03 – 현재 · LLM / Agent / 생성형 AI · Full-stack
-
담당업무: Agentic AI 기반 검색·챗봇 시스템 연구개발 및 AI 모델 서빙·인프라 구성
- SIGMOID AI — 생성형 LLM과 Semantic Search를 결합한 RAG 기반 AI 검색·챗봇 시스템 설계·개발
- OpenSearch + Vector DB 인덱싱 및 RAG 파이프라인 구축
- 검색 엔진 마이그레이션: Apache 2.0 라이선스 마지막 버전인 Elasticsearch 7.10.2에서 최신 OpenSearch 3.x로 전환 — 인덱스·매핑 이관, 쿼리·클라이언트 호환성 대응
- 범용 파일 인덱싱 API 개발 (PDF, HWP, PPT 등)
- LangChain / LangGraph 기반 AI 에이전트 설계 — 고정된 워크플로(predefined path)가 아닌, 모델이 자율 판단으로 툴 호출을 이어가는 동적(agentic) 오케스트레이션 구현
- Multi-Agent + 페르소나 기반 역할 분리로 응답 일관성·역할 유지
- 배포 후 지속 모니터링으로 검색 품질 측정, 시간 경과에 따른 성능 저하(drift) 탐지·대응
- 성과: 고객사 DB 연결 없이 범용 파일 인덱싱 실현, 다국어 입출력 지원, 최근 3개월 평균 3.8k 사용자
- SIGMOID LLM Manager — 도메인·상태 기반 AI 모델 자동 라우팅 LLM Gateway
- FastAPI 기반 API 인증 서버 및 모델 자동 라우팅 구현 — 모든 LLM 트래픽이 거쳐 가는 단일 게이트웨이(API 키 인증, IP 접근 제어, OpenAI 호환 프록시)
- Prometheus + Grafana 기반 모니터링·시각화
- 성과: 모델 상태(응답 지연·오류율) 기반 Failover 및 우선순위 라우팅, 동시 요청 100건 0.1초 내 처리, 월 평균 3만 건 API
- 공통 인프라 · 운영 — 사내 서버망과 클라우드 망을 함께 관리하며 서비스 인프라 구축·운영
- 네트워크: 서버 간 네트워크 구성·연결, 방화벽·ACL 정책 관리, 도메인/DNS 관리, SSL/TLS 인증서 발급·갱신
- 배포: Docker / docker-compose 기반 다중 서비스 구성·배포, Nginx 웹서버·로드밸런싱 구성으로 무중단 배포, GitLab CI · Jenkins 기반 CI/CD 파이프라인 구축, GPU 서버 세팅(드라이버·CUDA·NVIDIA Container Toolkit)
- 운영·관측: 로그 수집·관리, 장애 알림 룰 구성, 리소스 모니터링 및 용량 산정, DB 백업·복구
- 클라우드: 네이버 클라우드(NCP) 리소스(VPC·서버·스토리지) 관리 및 온프렘↔클라우드 연동, CSAP·GS 인증 취득 지원, KMS v2 키 관리
- AI 서비스: vLLM 기반 로컬 LLM·Reranker 추론 서빙 인프라 구축·운영 — H100 ×1 1대 + RTX 4090 ×2 3대(총 4대) GPU 서버에 모델 서빙 구성, 외부 API 없이 자체 서빙
- AI 가속기 도입 벤치마크 (단독 담당): Pliops SSD 레벨 KV 캐시 오프로딩 가속기에 대해 국내 유통사의 수입 검토를 위한 성능 검증을 전담 — vLLM 추론 서빙 환경에서 TTFT·프리필 지연, 처리량, KV 캐시 히트율·용량을 측정하는 벤치마크 시나리오 설계·수행 및 결과 리포팅
- 폐쇄망 로컬 LLM 서빙 검증: 폐쇄망 환경 도입을 위한 로컬 LLM 서빙 구축 — RTX 4090 ×1 / ×2, MacBook Pro 48GB, DGX Spark 하드웨어별 벤치마크 및 실사용 테스트로 도입 구성 비교·검증, Linux · CUDA 기반 AI 인프라 세팅 포함
슈어소프트테크 — AI기술팀 인턴
- 2024.07 – 2024.11 · AI / NLP / sLLM / Agent
-
담당업무: 도메인 특화 sLLM 파인튜닝 · LLMOps · 서빙 및 LLM 평가 벤치마크 연구개발
- Code2Chart — 도메인 특화 sLLM 파인튜닝
- 데이터 정제·수집, 학습 파이프라인 구축, Docker 기반 sgLang 서빙
- LLMOps: MLflow · WandB · DVC로 튜닝 데이터셋 · 하이퍼파라미터 레시피 · 모델 아티팩트를 버전 관리 — 실험 이력 추적 및 재현 가능한 학습 파이프라인 구성
- SFT 기법(FFT · LoRA · QLoRA) 비교 및 양자화(Quantization) 기반 경량 파인튜닝, CoT 데이터 학습
- H100 80GB × 8 단일 노드에서 FSDP · DeepSpeed ZeRO 샤딩으로 분산 학습 환경 구성 — LLaMA-Factory 기반 학습 구성에 Mixed Precision · Gradient Checkpointing · FlashAttention · FlashInfer 커널 백엔드 적용
- vLLM Prefill/Decode 분리(P/D Disaggregation) 서빙 구성 — KV 캐시 전송 경로 및 인스턴스별 병렬도·배치 정책 분리 조정
- 성과: Base 모델 33점 → 파인튜닝 67점 (2배), 사내 배포 및 QA 실사용 검증
- sLLM Benchmark — 도메인 특화 LLM-as-a-Judge 벤치마크
- Chain-of-Thought(CoT) 기반 평가 템플릿, 10개 세부 항목 설계
- 성과: GT 데이터 100개 기준 신뢰도 89.5%
- 기타: LangChain 기반 태스크 자동화 에이전트 프로토타입 개발
AI 활용 역량
- Claude Code 기반 업무 자동화 — AI 코딩 에이전트를 실무 전반에 도입해 반복 작업을 자동화하고 개발 생산성을 대폭 향상 (Claude Code 메인 + Codex 서브 멀티모델을 Orca ADE(Agentic Development Environment) 환경에서 활용)
- 커스텀 스킬 개발·공유: 사내에서 자주 쓰는 업무를 스킬(Skill)로 직접 제작해 팀에 공유·활용
- 에이전트 하네스 구축: oh-my-claudecode 등 플러그인·스킬·서브에이전트를 조합해 멀티에이전트 오케스트레이션 환경 구성 — 구현 → QA · 테스트 · 검증까지 자동으로 수행되는 파이프라인 운영
- Hooks 기반 정책 자동화: 훅(Hooks)으로 개발 정책·점검 규칙을 자동 강제
- 토큰 최적화: 토큰 최적화 프록시(RTK) 를 실무에 적용해 토큰 사용량·비용 절감
- 컨텍스트 엔지니어링: CLAUDE.md로 프로젝트 공용 지침·개발 방식·점검 방식을 정의해 에이전트의 정확도와 일관성 확보
- AI 활용 Full-stack 확장: 백엔드·AI 중심 업무를 AI 코딩 에이전트를 활용해 프론트엔드까지 Full-stack으로 확장, 서비스 개발 전 영역 겸임
- 사내 AI 확산 주도 — 기존 개발 지식(인프라·CI/CD·테스트)과 AI 에이전트를 결합한 노하우를 조직에 전파
- 이노디스AX에서 Claude Code 활용 사내 세미나 개최 — 토큰 최적화 프록시, codebase-mcp 등 실전 활용법 공유
- 프로젝트별 표준 가이드·설정 구축: 프로젝트마다 CLAUDE.md·공용 설정·개발 표준 가이드를 정비해 팀 단위 AI 도입 기반 마련
- 비개발 직군 대상 AX 교육 — 현업 부서가 업무에 AI를 활용할 수 있도록 교육·지원
프로젝트
| 프로젝트 | 분야 |
|---|---|
| GuardianWatch | 비전 AI 기반 아동 안전 모니터링 (졸업작품). CCTV 실시간 다중 객체 추적(MOT)과 행동 분석으로 위험 상황을 감지하고, 아이의 위치·행동·위험 알림을 부모 앱에 실시간 전송 |
| resumable-upload | 의존성 없이 Python으로 구현한 TUS 재개 가능 업로드 프로토콜 (v1.0.0) — 서버 & 클라이언트 |
| Distributed HAR System for Scalable Wi-Fi Sensing | 프라이버시 보존 인간 행동 인식 (INC Lab, 2023.03–2024.08). 라즈베리파이 4B + Nexmon CSI 기반 분산 수집 노드, Wi-Fi CSI 신호를 스펙트로그램 이미지로 변환해 ViT로 분류 — 데이터의 1/N만으로 정확도 90%, 도청 모델 정확도는 ~30% 이하로 억제 |
오픈소스
| 프로젝트 | 기여 내용 |
|---|---|
| vllm-project/vllm | GGUF 모델 로딩 경로 개선 및 오프라인·클라우드 스토리지 로딩 버그 수정 — 고성능 LLM 추론 엔진 #39471 · #39155 · #29582 · #29137 |
| vllm-project/vllm-metal | Metal paged 경로의 구조화 출력(grammar / structured output) 지원 구현 — Apple Silicon용 vLLM 하드웨어 플러그인 #280 |
| hiyouga/LLaMA-Factory | 학습 설정 로딩을 OmegaConf 기반으로 개선 — 통합 LLM 파인튜닝 프레임워크 #8505 |
| tus/tus.io | 직접 구현한 resumable-upload를 공식 구현체 목록에 등재 — TUS 재개 가능 업로드 프로토콜 공식 사이트 #603 · #545 |
| vibrantlabsai/ragas | 설치 및 퀵스타트 문서 오류 수정 — LLM 애플리케이션 평가 프레임워크 #2564 · #2563 |
| langflow-ai/langflow | 한국어 README 추가 — 비주얼 LLM 워크플로우 빌더 #2883 |
| OpenBMB/ChatDev | GPT-4o / GPT-4o-mini 모델 지원 추가, 한국어 README 보완 — 멀티에이전트 소프트웨어 개발 #405 · #404 |
| OpenBMB/RepoAgent | RAG 모듈의 시스템 프롬프트 오류 수정 — LLM 기반 저장소 문서화 #74 |
| stitionai/devika | GPT-4o-mini 모델 지원 추가 — 에이전트형 AI 소프트웨어 엔지니어 #630 |
| context-labs/autodoc | GPT-4o / GPT-4o-mini 모델 지원 추가 — LLM 기반 코드베이스 문서화 #46 |
| pqrs-org/KE-complex_modifications | 기존 단축키를 유지하는 원화(₩) 입력 키 매핑 규칙 추가 — Karabiner-Elements 개발자 도구 #1809 |
연구 및 활동
- Pseudo Lab
- Mammogram Processing for All — 맘모그램 영상 분석용 오픈소스 뷰어/편집 프로그램 및 관련 문서 제공
- INC Lab (무선통신), 가천대학교 (2023.03 – 2024.08)
- WiFi CSI 기반 인간 행동 인식, 프라이버시 보존 분산 센싱
- 라즈베리파이 4B + Nexmon CSI로 CSI 수집 노드를 구성해 분산 센싱 환경 구축 및 데이터 수집
- 논문: Distributed HAR System for Scalable Wi-Fi Sensing — 확장 가능한 WiFi 센싱을 위한 분산 행동 인식 시스템 (CSI 스펙트로그램 → ViT 분류, 데이터 1/N로 정확도 90%, 도청 모델 정확도 ~30% 이하)
- BDA (Big Data Association, 現 BDAI) (2024.03 – 2024.08)
- EDA, 통계 분석, NLP 전처리
자격증 및 수상
- 정보처리기사 (2024.09)
- 데이터분석준전문가 (ADsP) (2024.06)
- SQL개발자 (SQLD) (2024.04)
- TOEIC Speaking 150 — Intermediate High (2024.08)
- 🥈 K-디지털 챌린지: NET Challenge Camp S10 은상 (과학기술정보통신부, 2023.12)
Education
- M.S. in Department of Big Data, Sungkyunkwan University (2026.03 – 2028.02, expected)
- B.S. in Department of Software, School of AI·Software, Gachon University (2019.03 – 2025.02)
- GPA 3.81 / 4.5
Experience
innodisAX — Associate Research Engineer, Development Team
- 2025.03 – Present · LLM / Agent / Generative AI · Full-stack
-
Scope: R&D on agentic-AI search and chatbot systems, and AI model serving & infrastructure
- SIGMOID AI — Designed and built a RAG-based AI search and chatbot system combining generative LLMs with semantic search
- Built OpenSearch + Vector DB indexing and a RAG pipeline
- Search engine migration: moved from Elasticsearch 7.10.2, the last Apache-2.0-licensed release, to the latest OpenSearch 3.x — index/mapping migration and query/client compatibility work
- Developed a general-purpose file indexing API (PDF, HWP, PPT, etc.)
- Designed AI agents with LangChain / LangGraph — dynamic, agentic orchestration where the model autonomously directs its own tool calls, rather than a predefined workflow path
- Role separation via multi-agent + persona design for consistent responses and stable role adherence
- Continuous post-deployment monitoring to measure search quality and counter performance drift over time
- Impact: general-purpose file indexing without connecting to customer DBs, multilingual I/O, avg. 3.8k users over the last 3 months
- SIGMOID LLM Manager — an LLM Gateway for domain- and status-based automatic AI model routing
- Implemented a FastAPI-based API auth server and automatic model routing — a single gateway all LLM traffic passes through (API key auth, IP access control, OpenAI-compatible proxy)
- Monitoring and visualization with Prometheus + Grafana
- Impact: failover and priority routing based on model status (latency, error rate), 100 concurrent requests handled within 0.1s, ~30k API calls/month
- Shared infrastructure & operations — managed both the on-premise server network and the cloud network, building and running the service infrastructure
- Networking: server-to-server network setup, firewall / ACL policy management, domain & DNS management, SSL/TLS certificate issuance and renewal
- Deployment: multi-service setup and deployment with Docker / docker-compose, zero-downtime deployment via Nginx web servers and load balancing, CI/CD pipelines with GitLab CI / Jenkins, GPU server provisioning (drivers, CUDA, NVIDIA Container Toolkit)
- Operations & observability: log collection and retention, alerting rules for incidents, resource monitoring and capacity planning, DB backup and recovery
- Cloud: Naver Cloud (NCP) resource management (VPC, servers, storage) and on-prem ↔ cloud integration, supported CSAP / GS certification, KMS v2 key management
- AI services: built and operated the vLLM-based local LLM / Reranker inference serving infrastructure — model serving across four GPU servers — one with 1× H100 and three with 2× RTX 4090, served in-house without external APIs
- AI accelerator benchmarking (sole owner): led the performance validation for a Korean distributor’s import evaluation of Pliops SSD-level KV-cache offload accelerators — designed and ran the benchmark scenarios on vLLM inference serving (TTFT / prefill latency, throughput, KV-cache hit rate and capacity) and reported the results
- Air-gapped local LLM serving validation: built local LLM serving for closed-network deployment — benchmarked and real-use-tested candidate setups across RTX 4090 ×1 / ×2, MacBook Pro 48GB, and DGX Spark hardware, including Linux / CUDA-based AI infrastructure provisioning
Suresofttech — AI Technology Team Intern
- 2024.07 – 2024.11 · AI / NLP / sLLM / Agent
-
Scope: Domain-specific sLLM fine-tuning, LLMOps, serving, and LLM evaluation benchmark R&D
- Code2Chart — domain-specific sLLM fine-tuning
- Data cleaning/collection, training pipeline, Docker-based sgLang serving
- LLMOps: versioned tuning datasets, hyperparameter recipes, and model artifacts with MLflow · WandB · DVC — experiment tracking and a reproducible training pipeline
- Compared SFT methods (FFT · LoRA · QLoRA) and applied quantization for lightweight fine-tuning, trained on Chain-of-Thought (CoT) data
- Distributed training on a single node with 8× H100 80GB — FSDP / DeepSpeed ZeRO sharding, with mixed precision, gradient checkpointing, and FlashAttention / FlashInfer kernel backends applied in a LLaMA-Factory training setup
- vLLM prefill/decode disaggregated (P/D) serving — KV-cache transfer path and per-instance parallelism/batching policies tuned separately
- Impact: base model 33 → fine-tuned 67 (2×), internal deployment and QA real-use validation
- sLLM Benchmark — domain-specific LLM-as-a-Judge benchmark
- Chain-of-Thought (CoT) evaluation templates, 10 detailed criteria
- Impact: 89.5% reliability on 100 ground-truth samples
- Other: prototyped a LangChain-based task automation agent
AI-Assisted Development
- Claude Code-driven workflow automation — adopted AI coding agents across day-to-day work, automating repetitive tasks and significantly boosting development productivity (Claude Code as the primary agent and Codex as a secondary model, inside the Orca ADE — Agentic Development Environment)
- Custom skill development & sharing: built frequently-used internal workflows into reusable Skills and shared them across the team
- Agent harness engineering: composed plugins, skills, and subagents (e.g. oh-my-claudecode) into a multi-agent orchestration setup — a pipeline that runs implementation through QA, testing, and verification automatically
- Hook-based policy automation: enforced development and review policies automatically via hooks
- Token optimization: applied a token-optimizing proxy (RTK) in production use, cutting token usage and cost
- Context engineering: defined project-wide conventions, development practices, and review criteria in CLAUDE.md for accurate, consistent agent behavior
- AI-augmented full-stack: extended a backend/AI-focused role into full-stack development by leveraging AI coding agents, covering the frontend side of services as well
- Driving internal AI adoption — spread know-how combining existing engineering skills (infrastructure, CI/CD, testing) with AI agents across the organization
- Hosted an in-house Claude Code seminar at innodisAX — covering the token-optimizing proxy, codebase-mcp, and hands-on usage patterns
- Per-project standard guides & configuration: set up CLAUDE.md, shared settings, and development standards for each project, laying the groundwork for team-wide AI adoption
- AX training for non-developer roles — helped business teams apply AI to their daily work
Projects
| Project | Domain |
|---|---|
| GuardianWatch | vision-AI childcare safety monitoring (graduation project). Real-time multi-object tracking (MOT) and behavior analysis on CCTV to detect hazards and push children’s location, behavior, and danger alerts to a parent app |
| resumable-upload | Zero-dependency Python implementation of the TUS resumable upload protocol (v1.0.0) — server & client. |
| Distributed HAR System for Scalable Wi-Fi Sensing | privacy-preserving human activity recognition (INC Lab, 2023.03–2024.08). Distributed capture nodes on Raspberry Pi 4B + Nexmon CSI; Wi-Fi CSI signals turned into spectrogram images and classified with a ViT — 90% accuracy using only 1/N of the data; eavesdropper model accuracy kept below ~30% |
Open Source
| Project | Contribution |
|---|---|
| vllm-project/vllm | Improved the GGUF model loading path; fixed offline and cloud-storage loading bugs — High-throughput LLM inference engine #39471 · #39155 · #29582 · #29137 |
| vllm-project/vllm-metal | Implemented grammar / structured-output support for the Metal paged path — Hardware plugin for vLLM on Apple Silicon #280 |
| hiyouga/LLaMA-Factory | Switched training-config loading to OmegaConf — Unified LLM fine-tuning framework #8505 |
| tus/tus.io | Listed my own resumable-upload on the official implementations page — Official site for the TUS resumable upload protocol #603 · #545 |
| vibrantlabsai/ragas | Fixed errors in the install and quickstart documentation — Evaluation framework for LLM applications #2564 · #2563 |
| langflow-ai/langflow | Added a Korean README — Visual LLM workflow builder #2883 |
| OpenBMB/ChatDev | Added GPT-4o / GPT-4o-mini model support and improved the Korean README — Multi-agent software development #405 · #404 |
| OpenBMB/RepoAgent | Fixed a broken system prompt in the RAG module — LLM-powered repository documentation #74 |
| stitionai/devika | Added GPT-4o-mini model support — Agentic AI software engineer #630 |
| context-labs/autodoc | Added GPT-4o / GPT-4o-mini model support — LLM-based codebase documentation #46 |
| pqrs-org/KE-complex_modifications | Added a Korean-won key mapping rule that preserves the existing shortcut — Karabiner-Elements developer tooling #1809 |
Research & Activities
- Pseudo Lab
- Mammogram Processing for All — Open-source mammogram viewer/editor with accompanying documentation
- INC Lab (Wireless Communication), Gachon University (2023.03 – 2024.08)
- WiFi CSI-based human activity recognition, privacy-preserving distributed sensing
- Built the distributed sensing testbed and collected data with Raspberry Pi 4B + Nexmon CSI capture nodes
- Paper: Distributed HAR System for Scalable Wi-Fi Sensing — distributed activity recognition for scalable Wi-Fi sensing (CSI spectrograms classified by a ViT, 90% accuracy with 1/N of the data, eavesdropper accuracy kept below ~30%)
- BDA (Big Data Association, now BDAI) (2024.03 – 2024.08)
- EDA, statistical analysis, NLP preprocessing
Certifications & Awards
- Engineer Information Processing (2024.09)
- Advanced Data Analytics Semi-Professional (ADsP) (2024.06)
- SQL Developer (SQLD) (2024.04)
- TOEIC Speaking 150 — Intermediate High (2024.08)
- 🥈 K-Digital Challenge: NET Challenge Camp S10 — Silver Prize (Ministry of Science and ICT, 2023.12)
Skills
| 구분Category | 스킬Skills |
|---|---|
| LLM / NLP | LangChainLangGraphRAGRerankerLLM-as-a-JudgeAgent |
| Fine-tuning | SFTFFTLoRAQLoRAQuantizationLLaMA-Factory |
| LLM Serving | vLLMsgLangTGI |
| Backend | PythonFastAPIDjangoFlask |
| ML / DL | PyTorchTransformersCUDAComputer VisionOpenCV |
| MLOps | DVCWandBMLflow |
| Data / Search | ElasticsearchOpenSearchMySQLSQLPandasMATLAB |
| Infra | LinuxDockerGitNginxGitLab CIJenkinsGitHub ActionsNaver Cloud (NCP)PrometheusGrafana |