Curriculum Vitae (CV)

학력

  • 성균관대학교 빅데이터학과 석사 (2026.03 – 현재)
  • 가천대학교 AI·소프트웨어학부 소프트웨어전공 학사 (2019.03 – 2025.02)
    • 학점 3.81 / 4.5

경력

이노디스AX — 개발팀 주임연구원

  • 2025.03 – 현재 · NLP / 생성형 AI

  • SIGMOID AI — 생성형 LLM과 Semantic Search를 결합한 RAG 기반 AI 검색 시스템 설계·개발
    • OpenSearch + Vector DB 인덱싱 및 RAG 파이프라인 구축
    • 검색 엔진 마이그레이션: Apache 2.0 라이선스 마지막 버전인 Elasticsearch 7.10.2에서 최신 OpenSearch 3.x로 전환 — 인덱스·매핑 이관, 쿼리·클라이언트 호환성 대응
    • 범용 파일 인덱싱 API 개발 (PDF, HWP, PPT 등)
    • LangChain / LangGraph 기반 AI 에이전트 설계
    • Multi-Agent + 페르소나 기반 역할 분리로 응답 일관성·역할 유지
    • 배포 후 지속 모니터링으로 검색 품질 측정, 시간 경과에 따른 성능 저하(drift) 탐지·대응
    • 성과: 고객사 DB 연결 없이 범용 파일 인덱싱 실현, 다국어 입출력 지원, 최근 3개월 평균 3.8k 사용자
  • SIGMOID LLM Manager — 도메인·상태 기반 AI 모델 자동 라우팅 시스템
    • FastAPI 기반 API 인증 서버 및 모델 자동 라우팅 구현
    • Prometheus + Grafana 기반 모니터링·시각화
    • 성과: 모델 상태(응답 지연·오류율) 기반 Failover 및 우선순위 라우팅, 동시 요청 100건 0.1초 내 처리, 월 평균 3만 건 API
  • 공통 인프라 · 운영 — 사내 서버망과 클라우드 망을 함께 관리하며 서비스 인프라 구축·운영
    • 네트워크: 서버 간 네트워크 구성·연결, 방화벽·ACL 정책 관리, 도메인/DNS 관리, SSL/TLS 인증서 발급·갱신
    • 배포: Docker / docker-compose 기반 다중 서비스 구성·배포, Nginx 웹서버·로드밸런싱 구성으로 무중단 배포, GitLab CI · Jenkins 기반 CI/CD 파이프라인 구축, GPU 서버 세팅(드라이버·CUDA·NVIDIA Container Toolkit)
    • 운영·관측: 로그 수집·관리, 장애 알림 룰 구성, 리소스 모니터링 및 용량 산정, DB 백업·복구
    • 클라우드: 네이버 클라우드(NCP) 리소스(VPC·서버·스토리지) 관리 및 온프렘↔클라우드 연동, CSAP·GS 인증 취득 지원, KMS v2 키 관리
    • AI 서비스: vLLM 기반 Reranker·LLM 추론 서비스 운영·배포
    • AI 가속기 도입 벤치마크 (단독 담당): Pliops SSD 레벨 KV 캐시 오프로딩 가속기에 대해 국내 유통사의 수입 검토를 위한 성능 검증을 전담 — vLLM 추론 서빙 환경에서 TTFT·프리필 지연, 처리량, KV 캐시 히트율·용량을 측정하는 벤치마크 시나리오 설계·수행 및 결과 리포팅

슈어소프트테크 — AI기술팀 인턴

  • 2024.07 – 2024.11 · AI / NLP

  • Code2Chart — 도메인 특화 sLLM 파인튜닝
    • 데이터 정제·수집, 학습 파이프라인 구축, MLOps / DVC, Docker 기반 sgLang 서빙
    • SFT 기법(FFT · LoRA · QLoRA) 비교 및 양자화(Quantization) 기반 경량 파인튜닝, CoT 데이터 학습
    • H100 80GB × 8 단일 노드에서 FSDP · DeepSpeed ZeRO 샤딩으로 분산 학습 환경 구성 — LLaMA-Factory 기반 학습 구성에 Mixed Precision · Gradient Checkpointing · FlashAttention · FlashInfer 커널 백엔드 적용
    • vLLM Prefill/Decode 분리(P/D Disaggregation) 서빙 구성 — KV 캐시 전송 경로 및 인스턴스별 병렬도·배치 정책 분리 조정
    • 성과: Base 모델 33점 → 파인튜닝 67점 (2배), 사내 배포 및 QA 실사용 검증
  • sLLM Benchmark — 도메인 특화 LLM-as-a-Judge 벤치마크
    • Chain-of-Thought(CoT) 기반 평가 템플릿, 10개 세부 항목 설계
    • 성과: GT 데이터 100개 기준 신뢰도 89.5%
  • 기타: LangChain 기반 태스크 자동화 에이전트 프로토타입 개발

프로젝트

프로젝트 분야
GuardianWatch 비전 AI 기반 아동 안전 모니터링 (졸업작품). CCTV 실시간 다중 객체 추적(MOT)과 행동 분석으로 위험 상황을 감지하고, 아이의 위치·행동·위험 알림을 부모 앱에 실시간 전송
resumable-upload 의존성 없이 Python으로 구현한 TUS 재개 가능 업로드 프로토콜 (v1.0.0) — 서버 & 클라이언트
Distributed HAR System for Scalable Wi-Fi Sensing 프라이버시 보존 인간 행동 인식 (INC Lab, 2023.03–2024.08). 라즈베리파이 4B + Nexmon CSI 기반 분산 수집 노드, Wi-Fi CSI 신호를 스펙트로그램 이미지로 변환ViT로 분류 — 데이터의 1/N만으로 정확도 90%, 도청 모델 정확도는 ~30% 이하로 억제

오픈소스

기여 내역

프로젝트 기여 내용
vllm-project/vllm GGUF 모델 로딩 경로 개선 및 오프라인·클라우드 스토리지 로딩 버그 수정 — 고성능 LLM 추론 엔진 #39471 · #39155 · #29582 · #29137
vllm-project/vllm-metal Metal paged 경로의 구조화 출력(grammar / structured output) 지원 구현 — Apple Silicon용 vLLM 하드웨어 플러그인 #280
hiyouga/LLaMA-Factory 학습 설정 로딩을 OmegaConf 기반으로 개선 — 통합 LLM 파인튜닝 프레임워크 #8505
tus/tus.io 직접 구현한 resumable-upload를 공식 구현체 목록에 등재 — TUS 재개 가능 업로드 프로토콜 공식 사이트 #603 · #545
vibrantlabsai/ragas 설치 및 퀵스타트 문서 오류 수정 — LLM 애플리케이션 평가 프레임워크 #2564 · #2563
langflow-ai/langflow 한국어 README 추가 — 비주얼 LLM 워크플로우 빌더 #2883
OpenBMB/ChatDev GPT-4o / GPT-4o-mini 모델 지원 추가, 한국어 README 보완 — 멀티에이전트 소프트웨어 개발 #405 · #404
OpenBMB/RepoAgent RAG 모듈의 시스템 프롬프트 오류 수정 — LLM 기반 저장소 문서화 #74
stitionai/devika GPT-4o-mini 모델 지원 추가 — 에이전트형 AI 소프트웨어 엔지니어 #630
context-labs/autodoc GPT-4o / GPT-4o-mini 모델 지원 추가 — LLM 기반 코드베이스 문서화 #46
pqrs-org/KE-complex_modifications 기존 단축키를 유지하는 원화(₩) 입력 키 매핑 규칙 추가 — Karabiner-Elements 개발자 도구 #1809

연구 및 활동

  • Pseudo Lab
  • INC Lab (무선통신), 가천대학교 (2023.03 – 2024.08)
    • WiFi CSI 기반 인간 행동 인식, 프라이버시 보존 분산 센싱
    • 라즈베리파이 4B + Nexmon CSI로 CSI 수집 노드를 구성해 분산 센싱 환경 구축 및 데이터 수집
    • 논문: Distributed HAR System for Scalable Wi-Fi Sensing — 확장 가능한 WiFi 센싱을 위한 분산 행동 인식 시스템 (CSI 스펙트로그램 → ViT 분류, 데이터 1/N로 정확도 90%, 도청 모델 정확도 ~30% 이하)
  • BDA (Big Data Association, 現 BDAI) (2024.03 – 2024.08)
    • EDA, 통계 분석, NLP 전처리

자격증 및 수상

  • 정보처리기사 (2024.09)
  • 데이터분석준전문가 (ADsP) (2024.06)
  • SQL개발자 (SQLD) (2024.04)
  • TOEIC Speaking 150 — Intermediate High (2024.08)
  • 🥈 K-디지털 챌린지: NET Challenge Camp S10 은상 (과학기술정보통신부, 2023.12)

Education

  • M.S. in Department of Big Data, Sungkyunkwan University (2026.03 – Present)
  • B.S. in Department of Software, School of AI·Software, Gachon University (2019.03 – 2025.02)
    • GPA 3.81 / 4.5

Experience

innodisAX — Associate Research Engineer, Development Team

  • 2025.03 – Present · NLP / Generative AI

  • SIGMOID AI — Designed and built a RAG-based AI search system combining generative LLMs with semantic search
    • Built OpenSearch + Vector DB indexing and a RAG pipeline
    • Search engine migration: moved from Elasticsearch 7.10.2, the last Apache-2.0-licensed release, to the latest OpenSearch 3.x — index/mapping migration and query/client compatibility work
    • Developed a general-purpose file indexing API (PDF, HWP, PPT, etc.)
    • Designed AI agents with LangChain / LangGraph
    • Role separation via multi-agent + persona design for consistent responses and stable role adherence
    • Continuous post-deployment monitoring to measure search quality and counter performance drift over time
    • Impact: general-purpose file indexing without connecting to customer DBs, multilingual I/O, avg. 3.8k users over the last 3 months
  • SIGMOID LLM Manager — domain- and status-based automatic AI model routing system
    • Implemented a FastAPI-based API auth server and automatic model routing
    • Monitoring and visualization with Prometheus + Grafana
    • Impact: failover and priority routing based on model status (latency, error rate), 100 concurrent requests handled within 0.1s, ~30k API calls/month
  • Shared infrastructure & operations — managed both the on-premise server network and the cloud network, building and running the service infrastructure
    • Networking: server-to-server network setup, firewall / ACL policy management, domain & DNS management, SSL/TLS certificate issuance and renewal
    • Deployment: multi-service setup and deployment with Docker / docker-compose, zero-downtime deployment via Nginx web servers and load balancing, CI/CD pipelines with GitLab CI / Jenkins, GPU server provisioning (drivers, CUDA, NVIDIA Container Toolkit)
    • Operations & observability: log collection and retention, alerting rules for incidents, resource monitoring and capacity planning, DB backup and recovery
    • Cloud: Naver Cloud (NCP) resource management (VPC, servers, storage) and on-prem ↔ cloud integration, supported CSAP / GS certification, KMS v2 key management
    • AI services: operated and deployed vLLM-based Reranker / LLM inference services
    • AI accelerator benchmarking (sole owner): led the performance validation for a Korean distributor’s import evaluation of Pliops SSD-level KV-cache offload accelerators — designed and ran the benchmark scenarios on vLLM inference serving (TTFT / prefill latency, throughput, KV-cache hit rate and capacity) and reported the results

Suresofttech — AI Technology Team Intern

  • 2024.07 – 2024.11 · AI / NLP

  • Code2Chart — domain-specific sLLM fine-tuning
    • Data cleaning/collection, training pipeline, MLOps / DVC, Docker-based sgLang serving
    • Compared SFT methods (FFT · LoRA · QLoRA) and applied quantization for lightweight fine-tuning, trained on Chain-of-Thought (CoT) data
    • Distributed training on a single node with 8× H100 80GBFSDP / DeepSpeed ZeRO sharding, with mixed precision, gradient checkpointing, and FlashAttention / FlashInfer kernel backends applied in a LLaMA-Factory training setup
    • vLLM prefill/decode disaggregated (P/D) serving — KV-cache transfer path and per-instance parallelism/batching policies tuned separately
    • Impact: base model 33 → fine-tuned 67 (2×), internal deployment and QA real-use validation
  • sLLM Benchmark — domain-specific LLM-as-a-Judge benchmark
    • Chain-of-Thought (CoT) evaluation templates, 10 detailed criteria
    • Impact: 89.5% reliability on 100 ground-truth samples
  • Other: prototyped a LangChain-based task automation agent

Projects

Project Domain
GuardianWatch vision-AI childcare safety monitoring (graduation project). Real-time multi-object tracking (MOT) and behavior analysis on CCTV to detect hazards and push children’s location, behavior, and danger alerts to a parent app
resumable-upload Zero-dependency Python implementation of the TUS resumable upload protocol (v1.0.0) — server & client.
Distributed HAR System for Scalable Wi-Fi Sensing privacy-preserving human activity recognition (INC Lab, 2023.03–2024.08). Distributed capture nodes on Raspberry Pi 4B + Nexmon CSI; Wi-Fi CSI signals turned into spectrogram images and classified with a ViT90% accuracy using only 1/N of the data; eavesdropper model accuracy kept below ~30%

Open Source

Contributions

Project Contribution
vllm-project/vllm Improved the GGUF model loading path; fixed offline and cloud-storage loading bugs — High-throughput LLM inference engine #39471 · #39155 · #29582 · #29137
vllm-project/vllm-metal Implemented grammar / structured-output support for the Metal paged path — Hardware plugin for vLLM on Apple Silicon #280
hiyouga/LLaMA-Factory Switched training-config loading to OmegaConf — Unified LLM fine-tuning framework #8505
tus/tus.io Listed my own resumable-upload on the official implementations page — Official site for the TUS resumable upload protocol #603 · #545
vibrantlabsai/ragas Fixed errors in the install and quickstart documentation — Evaluation framework for LLM applications #2564 · #2563
langflow-ai/langflow Added a Korean README — Visual LLM workflow builder #2883
OpenBMB/ChatDev Added GPT-4o / GPT-4o-mini model support and improved the Korean README — Multi-agent software development #405 · #404
OpenBMB/RepoAgent Fixed a broken system prompt in the RAG module — LLM-powered repository documentation #74
stitionai/devika Added GPT-4o-mini model support — Agentic AI software engineer #630
context-labs/autodoc Added GPT-4o / GPT-4o-mini model support — LLM-based codebase documentation #46
pqrs-org/KE-complex_modifications Added a Korean-won key mapping rule that preserves the existing shortcut — Karabiner-Elements developer tooling #1809

Research & Activities

  • Pseudo Lab
  • INC Lab (Wireless Communication), Gachon University (2023.03 – 2024.08)
    • WiFi CSI-based human activity recognition, privacy-preserving distributed sensing
    • Built the distributed sensing testbed and collected data with Raspberry Pi 4B + Nexmon CSI capture nodes
    • Paper: Distributed HAR System for Scalable Wi-Fi Sensing — distributed activity recognition for scalable Wi-Fi sensing (CSI spectrograms classified by a ViT, 90% accuracy with 1/N of the data, eavesdropper accuracy kept below ~30%)
  • BDA (Big Data Association, now BDAI) (2024.03 – 2024.08)
    • EDA, statistical analysis, NLP preprocessing

Certifications & Awards

  • Engineer Information Processing (2024.09)
  • Advanced Data Analytics Semi-Professional (ADsP) (2024.06)
  • SQL Developer (SQLD) (2024.04)
  • TOEIC Speaking 150 — Intermediate High (2024.08)
  • 🥈 K-Digital Challenge: NET Challenge Camp S10 — Silver Prize (Ministry of Science and ICT, 2023.12)

Skills

구분Category 스킬Skills
LLM / NLP LangChainLangGraphRAGRerankerLLM-as-a-JudgeAgent
Fine-tuning SFTFFTLoRAQLoRAQuantizationLLaMA-Factory
LLM Serving vLLMsgLangTGI
Backend PythonFastAPIDjangoFlask
ML / DL PyTorchTransformersCUDAComputer VisionOpenCV
MLOps DVCWandBMLflow
Data / Search ElasticsearchOpenSearchMySQLSQLPandasMATLAB
Infra LinuxDockerGitNginxGitLab CIJenkinsGitHub ActionsNaver Cloud (NCP)PrometheusGrafana