Computer Vision R&D Roadmap: CNN, GAN, and 3D Reconstruction

A structured study guide covering core CV algorithms — from R-CNN and YOLO to GAN variants and deep learning-based 3D reconstruction techniques

Captain Ethan
Captain Paul
Maritime 4.0 · AI, Data & Cyber Security  ·  linkedin.com/in/shipjobs

Computer Vision (CV) is one of the most rapidly advancing fields in AI, enabling machines to interpret and understand visual data from the world. This R&D roadmap covers the core algorithmic pillars — from CNN-based object detection (R-CNN, YOLO) to generative models (GAN variants) and deep learning-based 3D reconstruction — along with a curated list of R&D topics and practical study resources.

Key Terms
CV — Computer Vision
CNN — Convolutional Neural Network
R-CNN — Region-based CNN
YOLO — You Only Look Once
GAN — Generative Adversarial Network
cGAN — Conditional GAN
DCGAN — Deep Convolutional GAN
SfM — Structure from Motion
RGB — Red Green Blue color model
YUV — Luminance + Chrominance color model

Ⅰ. Overview

Computer Vision Pipeline
📷 Input Image
⚙️ Preprocessing
🔍 Feature Extraction (CNN)
🎯 Detection / Classification
📊 Output
👁️
Computer Vision
Enabling machines to interpret and understand visual information from images and video — object detection, tracking, recognition, and generation.
🖼️
Image Processing
Fundamental operations for transforming and analyzing digital images — filtering, enhancement, segmentation, and feature extraction.

Ⅱ. Background — RGB vs YUV

RGB Model
Represents color as a combination of Red, Green, Blue channels. Directly corresponds to display hardware. Most common format for image capture and rendering.
YUV Model
Separates luminance (Y) from chrominance (U, V). More efficient for compression and human visual system modeling — widely used in video codecs (H.264, H.265).

Ⅲ. Core Algorithms

📦 CNN — Convolutional Neural Networks
R-CNN CVPR 2014
Region proposals + CNN features. Accurate but slow — processes each region independently.
💡 핵심 원리: Selective Search로 ~2000개 region proposal 추출 → 각 영역을 CNN으로 개별 처리 → SVM 분류. 정확하지만 영역별 CNN 순전파가 반복되어 이미지당 47초 소요. CNN 기반 물체 탐지의 출발점입니다.
Fast R-CNN ICCV 2015
Shared CNN computation via RoI pooling. ~9× faster than R-CNN, end-to-end training.
💡 핵심 원리: 이미지 전체를 CNN에 한 번만 통과시켜 feature map 생성. RoI Pooling으로 각 proposal 영역의 feature를 고정 크기로 추출 → FC → 분류+박스 회귀를 동시 학습. R-CNN 대비 9배 빠르고 end-to-end 학습 가능합니다.
Faster R-CNN NIPS 2015
Region Proposal Network (RPN) integrated into CNN. Near real-time detection, state-of-the-art accuracy.
💡 핵심 원리: Selective Search를 RPN으로 대체. RPN은 feature map을 공유하여 anchor box마다 objectness score + bbox 회귀를 예측. Proposal 생성이 GPU 위에서 10ms 내 처리됩니다. 5 fps로 당시 실시간에 근접.
⚡ YOLO — You Only Look Once

Single-pass detection: divides image into grid cells, each predicting bounding boxes and class probabilities simultaneously. Real-time speed with competitive accuracy. Foundation for YOLOv2–v8 variants widely used in maritime vessel detection.

🎨 GAN — Generative Adversarial Networks
cGAN (Conditional GAN) 2014
Generation conditioned on class labels or auxiliary information. Enables targeted output control.
💡 핵심 원리: Generator G(z|y)와 Discriminator D(x|y) 모두에 조건 정보 y를 입력. y는 클래스 레이블, 텍스트, 다른 이미지 등 어떤 형태든 가능. Pix2Pix(2017)의 이미지-이미지 변환 기반이 됩니다.
DCGAN ICLR 2016
Deep Convolutional GAN — stable training via strided convolutions, BatchNorm, and LeakyReLU.
💡 핵심 원리: FC Layer 제거, Fractional-strided Conv으로 업샘플링, BatchNorm + LeakyReLU 조합으로 학습 안정화. GAN 훈련의 표준 레시피를 확립했으며 학습된 잠재 공간(latent space)이 의미 있는 방향(벡터 연산)을 가짐을 최초로 시연.
CycleGAN ICCV 2017
Unpaired image-to-image translation using cycle-consistency loss. No paired training data required.
💡 핵심 원리: G: X→Y, F: Y→X 두 Generator + 두 Discriminator. Cycle-consistency: F(G(x)) ≈ x 로 쌍 없이도 도메인 변환 학습. 말→얼룩말, 여름→겨울 등 스타일 변환에 광범위하게 사용됩니다.
StyleGAN CVPR 2019
Style-based generator with AdaIN. Fine-grained control over image style at multiple scales.
💡 핵심 원리: Mapping Network z→w로 스타일 벡터 생성, AdaIN으로 각 레이어에 스타일 주입. 저해상도(구조)~고해상도(텍스처) 레이어에 서로 다른 스타일을 분리 제어 가능. StyleGAN2(2020)는 아티팩트 제거, StyleGAN3(2021)는 이미지 변환 등변성 개선.
🔷 3D Reconstruction — Generation-by-Generation
Generation Papers Characteristic
1st Gen MC-CNN [CVPR15]
Deep Embed [ICCV15]
Content-CNN [CVPR16]
Early CNN-based stereo matching — learning local patch similarity for disparity estimation
2nd Gen DispNet [CVPR16]
DeMoN [CVPR17]
End-to-end networks for disparity and depth+motion estimation from image pairs
3rd Gen DPSNet [ICLR19] Plane-sweep stereo with cost volume — multi-view depth estimation with learned cost aggregation

Ⅳ. R&D Topics

📸 Image Classification
Assign category labels to entire images using CNN architectures (ResNet, VGG, EfficientNet)
🎯 Object Detection
Locate and classify multiple objects with bounding boxes using R-CNN, YOLO, SSD
🔍 Object Tracking
Follow object positions across video frames using temporal features (SORT, DeepSORT)
🔷 3D Reconstruction
Recover 3D structure from 2D images using stereo, monocular depth, or NeRF
🏃 Action Recognition
Classify human activities from video sequences using 3D CNN or Transformer models
🔆 Image Super-Resolution
Enhance low-resolution images using GAN/CNN (SRGAN, ESRGAN) for detail recovery
✂️ Image Segmentation
Assign class labels per pixel — semantic, instance, and panoptic segmentation
👤 Human Face Analysis
Detection, recognition, landmark localization, and expression analysis from facial images
🚶 Person Re-ID
Match person identities across different cameras and viewpoints using metric learning
🎨 GAN for Vision
Generate, edit, and augment images via adversarial training for data synthesis
🖌️ Generative Models
VAE, Diffusion, Flow-based models for controllable high-quality image synthesis
🏙️ Scene Understanding
Holistic parsing of scenes including depth estimation, layout prediction, and semantic labeling
🔎 Image Retrieval
Find visually similar images using feature embedding and metric learning (ArcFace, triplet loss)
📝 CV + NLP
Visual question answering, image captioning, and vision-language models (CLIP, BLIP)
💡 Lightweight Architectures
MobileNet, ShuffleNet, SqueezeNet for edge deployment with limited compute
🔄 Domain Adaptation
Transfer models across different visual domains without re-labeling (UDA, fine-tuning)
🧠 Unsupervised Feature Learning
Learn visual representations without labels using contrastive learning (SimCLR, DINO, MAE)

Ⅴ. 실습 (Practice)

Recommended study resources and hands-on materials:

Captain's Take — CV as a Maritime Intelligence Layer

Computer Vision is no longer just an academic research topic — it is rapidly becoming a core intelligence layer in Maritime 4.0:

Object detection (YOLO, Faster R-CNN) enables automated vessel identification and collision avoidance in smart ship systems
GAN-based generation supports synthetic training data for scenarios that are rare or dangerous to capture in maritime environments
3D Reconstruction enables digital twin construction of port infrastructure and vessel hulls for predictive maintenance

Mastering these algorithms forms the research foundation for building next-generation maritime AI systems.

#ComputerVision #CNN #YOLO #GAN #3DReconstruction #DeepLearning #ImageProcessing #Maritime4.0
2020 →

졸업 이후 — CV의 새로운 계보

CNN·GAN 중심에서 Transformer, Diffusion Model, NeRF/3DGS로 패러다임이 전환됐습니다. 2020년 이후 Computer Vision은 역사상 가장 빠른 속도로 진화하고 있습니다.

Wave 1

Transformer의 CV 침투 (2020~2022)

ViT Vision Transformer [ICLR 2021]
이미지를 16×16 패치로 분할 후 각 패치를 토큰으로 취급, Transformer Encoder로 처리. CNN 귀납 편향(locality, translation equivariance) 없이 대규모 데이터에서 CNN을 능가.
💡 핵심 원리: 이미지 → N개 패치 → Linear projection → CLS 토큰 추가 → Positional embedding → Transformer Encoder → 분류. JFT-300M 등 초대규모 사전학습 필요. DeiT(2021)는 데이터 효율 훈련 증류로 이 문제를 해결했습니다.
DETR Detection Transformer [ECCV 2020]
NMS(Non-Maximum Suppression)와 anchor box를 완전히 제거. Encoder-Decoder Transformer + bipartite matching으로 end-to-end 물체 탐지. 탐지 파이프라인의 혁신.
💡 핵심 원리: N개 object query가 Decoder를 통해 이미지 feature와 cross-attention. 헝가리안 알고리즘으로 예측-GT 매칭 후 손실 계산. 후처리 불필요. 수렴이 느린 단점은 Deformable DETR(2021)이 개선했습니다.
CLIP Contrastive Language-Image Pre-training [ICML 2021]
4억 쌍의 이미지-텍스트 대조 학습. 이미지 인코더와 텍스트 인코더를 동일 공간에 정렬. Zero-shot 분류, 이미지 검색, Stable Diffusion의 텍스트 조건부 핵심.
💡 핵심 원리: 배치 내 N개 이미지-텍스트 쌍 중 올바른 쌍은 유사도 최대화, 나머지 N²−N 쌍은 최소화. "A photo of a {class}" 프롬프트만으로 임의 클래스 Zero-shot 분류가 가능합니다.
Wave 2

Diffusion Model — GAN을 대체한 생성 패러다임 (2020~2023)

DDPM Denoising Diffusion Probabilistic Models [NeurIPS 2020]
노이즈를 순차적으로 추가(Forward)했다가 역방향으로 제거(Reverse)하는 반복 생성. GAN의 훈련 불안정성 없이 고품질 이미지 합성. Stable Diffusion·DALL-E의 이론적 토대.
💡 핵심 원리: Forward: xₜ = √ᾱₜx₀ + √(1-ᾱₜ)ε. Reverse: U-Net이 각 스텝에서 노이즈 ε를 예측하여 제거. DDIM(2021)은 결정론적 샘플링으로 수십 배 빠른 추론을 실현했습니다.
LDM Stable Diffusion (Latent Diffusion) [CVPR 2022]
픽셀 공간이 아닌 VAE 잠재 공간에서 Diffusion 수행. CLIP 텍스트 인코더로 "a ship on stormy sea" 같은 텍스트 조건부 생성. 대중적 AI 이미지 생성의 시대를 열었습니다.
💡 핵심 원리: VAE Encoder로 이미지를 4×H/8×W/8 잠재 코드로 압축 → 압축 공간에서 DDPM 수행(연산량 8배 절감) → VAE Decoder로 복원. Cross-attention으로 텍스트 조건을 U-Net 모든 레이어에 주입합니다.
Wave 3

NeRF & 3D Gaussian Splatting — 3D 재건의 혁명 (2020~2023)

NeRF Neural Radiance Fields [ECCV 2020]
다시점 이미지에서 암묵적 신경망 표현으로 3D 장면 복원. MLP가 (x,y,z,θ,φ) → (RGB, σ) 를 학습하여 novel view synthesis. 3D 재건의 패러다임을 완전히 바꿨습니다.
💡 핵심 원리: 광선(ray) 위 여러 샘플 포인트의 (RGB, 밀도σ)를 적분하여 픽셀 색상 예측. Positional Encoding으로 고주파 세부 표현. 단점: 장면당 수시간 학습. Instant-NGP(2022)는 해시 인코딩으로 수초로 단축.
3DGS 3D Gaussian Splatting [SIGGRAPH 2023]
3D 장면을 수백만 개의 3D Gaussian으로 표현. 암묵적 MLP 대신 명시적 Gaussian 파라미터로 실시간(30fps↑) 렌더링. NeRF의 속도 한계를 극복한 차세대 3D 표현.
💡 핵심 원리: 각 Gaussian은 위치(μ), 공분산(Σ), 투명도(α), 색상(SH 계수)로 구성. Splatting으로 이미지 평면에 투영 후 알파 합성. SFM 포인트 클라우드로 초기화 → 최적화. 드론 촬영 영상에서 선박 3D 복원에 직접 적용 가능합니다.
SAM Segment Anything Model [ICCV 2023]
10억 개 마스크로 사전학습한 범용 분할 모델. 클릭·박스·텍스트 프롬프트로 어떤 객체든 즉시 분할. Zero-shot 세그멘테이션의 기준점.
💡 핵심 원리: Image Encoder(ViT-H) + Prompt Encoder(점/박스/텍스트) + Mask Decoder(Transformer). SA-1B 데이터셋(1,100만 이미지, 11억 마스크)으로 학습. SAM2(2024)는 동영상까지 확장하여 선박 실시간 추적에 응용됩니다.
Computer Vision 발전 타임라인
~2019 R-CNN계열 · YOLO · GAN · 3D Cost Volume — 이 포스트의 내용
2020~21 ViT · DETR · CLIP · NeRF · DDPM — Transformer & 신경 표현의 부상
2022~23 Stable Diffusion · SAM · 3D Gaussian Splatting · YOLOv8
2024~현재 SAM2 · Video Diffusion · 4D Gaussian · Multimodal Vision-Language Models

📚 Related Papers & References

1
Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation (R-CNN)
Ross Girshick et al. · UC Berkeley · CVPR 2014
arxiv.org/abs/1311.2524
2
Fast R-CNN
Ross Girshick · Microsoft Research · ICCV 2015
arxiv.org/abs/1504.08083
3
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
Shaoqing Ren et al. · Microsoft Research · NIPS 2015
arxiv.org/abs/1506.01497
4
You Only Look Once: Unified, Real-Time Object Detection (YOLO)
Joseph Redmon et al. · University of Washington · CVPR 2016
arxiv.org/abs/1506.02640
5
Generative Adversarial Networks (GAN)
Ian Goodfellow et al. · Université de Montréal · NIPS 2014
arxiv.org/abs/1406.2661
6
Unpaired Image-to-Image Translation using Cycle-Consistent Adversarial Networks (CycleGAN)
Jun-Yan Zhu et al. · UC Berkeley · ICCV 2017
arxiv.org/abs/1703.10593
7
FlowNet: Learning Optical Flow with Convolutional Networks (DispNet basis)
Philipp Fischer et al. · University of Freiburg · ICCV 2015
arxiv.org/abs/1504.06852
8
DeMoN: Depth and Motion Network for Learning Monocular Stereo
Benjamin Ummenhofer et al. · University of Freiburg · CVPR 2017
arxiv.org/abs/1612.02401
Captain Ethan
Captain Paul · In Sung Lee
Maritime 4.0 · AI, Data & Cyber Security
🔗 LinkedIn · shipjobs  ·  Collaborator: Lew, Julius, Jin, Morgan, Yeon
🔬 R&D Research

⚓ Join the ShipPaulJobs Community

Join →
Share

Comments

Top Ranked · All Posts

Popular Posts