- Primary maintainer of an internal FastAPI resource backend for an in-house collection framework: 32 endpoints across 8 routers, with a MongoDB to PostgreSQL dual-stack rearchitecture.사내 수집 프레임워크의 내부 FastAPI 리소스 백엔드 주 메인테이너: 8개 라우터에 걸친 32개 엔드포인트, MongoDB → PostgreSQL 듀얼 스택 재설계.
- Co-maintainer of the in-house scraping framework; introduced a Job domain model that became the basis for company-wide monitoring with Grafana and custom data-quality CLIs.사내 스크래핑 프레임워크 공동 메인테이너. Job 도메인 모델을 도입해 Grafana와 자체 데이터 품질 CLI 기반 전사 모니터링의 토대를 만들었습니다.
- Sole owner of a 28-backend Kubernetes egress platform with drain-rotate automation, synthetic health checks, self-healing, and a daily Slack health report.28개 백엔드 Kubernetes 이그레스 플랫폼 단독 오너: drain-rotate 자동화, 합성 헬스 체크, 셀프 힐링, 매일 아침 Slack 헬스 리포트.
- Designed a vertical AI pipeline end to end for an industrial-security customer: Slack command to Argo Workflow, scraper, in-house LLM report, and Slack-thread delivery.산업 보안 고객을 위한 버티컬 AI 파이프라인을 엔드투엔드로 설계: Slack 명령 → Argo Workflow → 스크래퍼 → 사내 LLM 리포트 → Slack 스레드 전달.
Career profile커리어 프로필
A concise hiring profile for Sungju Kim, Data & AI Systems Engineer.Data & AI Systems Engineer 김성주의 간결한 채용용 프로필입니다.
Data Platform / Software Engineer — 5 years across Korea, France, and the US.데이터 플랫폼 / 소프트웨어 엔지니어 — 한국·프랑스·미국에서 5년.
I design and operate large-scale data collection pipelines and the cloud and Kubernetes infrastructure behind them. Most recently, I worked as a Data Platform / Software Engineer at S2W in Seongnam, Korea, where I was the primary maintainer of an internal FastAPI resource backend and the sole owner of egress, monitoring, and data-quality tooling. Before that, I worked as a remote freelance Data Engineer for Quantum Analytica in Boston on PySpark and Snowflake ETLs, and as a Python Developer on the Scraping team at Data Impact by NielsenIQ in Paris, moving from internship to CDI (French permanent contract).대규모 데이터 수집 파이프라인과 그 아래의 클라우드·Kubernetes 인프라를 설계하고 운영합니다. 가장 최근에는 한국 성남의 S2W에서 데이터 플랫폼 / 소프트웨어 엔지니어로 일하며 내부 FastAPI 리소스 백엔드의 주 메인테이너이자 이그레스·모니터링·데이터 품질 도구의 단독 오너였습니다. 그 전에는 미국 보스턴의 Quantum Analytica에서 원격 프리랜스 데이터 엔지니어로 PySpark·Snowflake ETL을 맡았고, 프랑스 파리의 Data Impact by NielsenIQ 스크래핑 팀에서 Python 개발자로 인턴십을 거쳐 CDI(프랑스 무기계약 정규직)로 전환했습니다.
- Led an 8-pipeline PySpark ETL framework for cannabis market analytics, writing Parquet + Zstd outputs to S3 and PostgreSQL.대마초 시장 분석용 8개 파이프라인 PySpark ETL 프레임워크를 주도, Parquet + Zstd 출력을 S3와 PostgreSQL에 적재.
- Led a 4-person real estate ETL project into production on Databricks, powering a commercial dashboard.4인 부동산 ETL 프로젝트를 Databricks에서 프로덕션까지 리드해 상용 대시보드에 공급.
- Built a Snowflake warehouse from multi-source ingestion, helping close a client contract.멀티 소스 수집으로 Snowflake 웨어하우스를 구축해 고객사 계약 성사에 기여.
- Converted from internship to full-time CDI (French permanent contract) on a 25-person scraping team serving global e-commerce data clients, including L'Oréal, Coca-Cola, and Unilever.L'Oréal, Coca-Cola, Unilever 등 글로벌 이커머스 데이터 고객을 지원하는 25인 스크래핑 팀에서 인턴십을 마치고 CDI(프랑스 무기계약 정규직)로 전환.
- Built and maintained spiders for 57 countries on a custom Scrapy framework: 553 web and 108 mobile spider directories.자체 Scrapy 프레임워크 위에서 57개국 스파이더를 개발·운영: 웹 553개, 모바일 108개 스파이더 디렉터리.
- Contributed Cloudflare/Akamai anti-bot middleware, Redis cookie management, multi-mode pipelines, and a 15-month InfluxDB/Grafana/Slack monitoring stack with ML-based anomaly alerts.Cloudflare/Akamai 안티봇 미들웨어, Redis 쿠키 관리, 멀티 모드 파이프라인, ML 기반 이상 알림을 갖춘 15개월의 InfluxDB/Grafana/Slack 모니터링 스택에 기여.
I work across the layers needed to ship small data and AI systems end to end.작은 데이터·AI 시스템을 엔드투엔드로 출시하는 데 필요한 레이어 전반을 다룹니다.
Production systems I owned end to end or led.엔드투엔드로 소유했거나 주도한 프로덕션 시스템입니다.
Vertical AI Pipeline for an Industrial-Security Customer (S2W)산업 보안 고객을 위한 버티컬 AI 파이프라인 (S2W)
Owned the flow end to end: Slack command → Argo Workflow → site-specific scrapers → in-house LLM field extraction and report draft → results posted back in the same Slack thread. Delivered as a Slack self-service workflow for the customer.Slack 명령 → Argo Workflow → 사이트별 스크래퍼 → 사내 LLM 필드 추출·리포트 초안 → 같은 Slack 스레드로 결과 회신까지 전체 플로우를 엔드투엔드로 소유했습니다. 고객이 Slack에서 직접 실행하는 셀프서비스 워크플로우로 제공됩니다.
Self-Healing Egress Platform (S2W)셀프 힐링 이그레스 플랫폼 (S2W)
Owned three stages: vendor selection and contract negotiation, a Kubernetes IP-rotation platform, and a Postgres event ledger with drain-rotate automation and synthetic CONNECT health checks. Ran 28 backends across 5 providers behind HAProxy.벤더 선정·계약 협상, Kubernetes IP 로테이션 플랫폼, drain-rotate 자동화와 합성 CONNECT 헬스 체크를 갖춘 Postgres 이벤트 레저까지 3단계를 소유했습니다. HAProxy 뒤에서 5개 공급자에 걸친 28개 백엔드를 운영했습니다.
Concurrency-Safe Resource Backend for a Collection Framework (S2W)수집 프레임워크를 위한 동시성 안전 리소스 백엔드 (S2W)
Primary maintainer of an internal FastAPI backend with 8 routers and 32 v2 endpoints across MongoDB and PostgreSQL. Added atomic account acquire/release, Playwright session merge, dynamic seed scheduling, and capture-to-Airflow handoff. Led the MongoDB to PostgreSQL relational-path migration and a v3.0.0 architecture refresh.MongoDB·PostgreSQL에 걸쳐 8개 라우터, 32개 v2 엔드포인트를 가진 내부 FastAPI 백엔드의 주 메인테이너. 원자적 계정 획득/반환, Playwright 세션 병합, 동적 시드 스케줄링, 캡처의 Airflow 전달을 추가했습니다. MongoDB → PostgreSQL 관계형 경로 마이그레이션과 v3.0.0 아키텍처 리프레시를 주도했습니다.
Cannabis Market ETL Framework (Quantum Analytica)대마초 시장 ETL 프레임워크 (Quantum Analytica)
Led an 8-ETL PySpark framework that wrote Parquet + Zstd outputs to S3 and PostgreSQL, plus a Snowflake warehouse build that helped close a client contract.Parquet + Zstd 출력을 S3와 PostgreSQL에 적재하는 8개 ETL PySpark 프레임워크를 주도했고, 고객사 계약 성사에 기여한 Snowflake 웨어하우스도 구축했습니다.
Global E-commerce Data Collection — 57 countries (Data Impact by NielsenIQ)글로벌 이커머스 데이터 수집 — 57개국 (Data Impact by NielsenIQ)
Built and maintained spiders across 553 web and 108 mobile retailer directories in 57 countries on a custom Scrapy framework. Contributed Cloudflare/Akamai anti-bot middleware, a 15-month InfluxDB/Grafana monitoring stack with ML-based anomaly alerts, and a Bitbucket Pipelines to ScrapingHub auto-deploy.자체 Scrapy 프레임워크에서 57개국, 웹 553개·모바일 108개 리테일러 디렉터리의 스파이더를 개발·운영했습니다. Cloudflare/Akamai 안티봇 미들웨어, ML 기반 이상 알림을 갖춘 15개월 InfluxDB/Grafana 모니터링 스택, Bitbucket Pipelines → ScrapingHub 자동 배포에 기여했습니다.
Best channel for hiring conversations.채용 관련 대화에 가장 좋은 채널입니다.
For formal applications, please use the resume submitted through the application channel.공식 지원 시에는 지원 채널을 통해 제출된 이력서를 기준으로 확인해 주세요.
View the public AI lab →공개 AI Lab 보기 →