Computer Vision

Single-Stage Multi-Person Pose Estimation with Distribution-Oriented Feature Correction

A single-stage multi-person human pose-estimation network with distribution-oriented feature correction for accurate keypoint localization. (M.S. thesis, 2026)

Design of 3D Hand Mesh Reconstruction from Monocular Image

Graph Convolutional Networks (GCNs) are well-suited for human action recognition using skeleton data, as they handle non-Euclidean structures like human joints and avoid issues with environmental noise affecting RGB images. However, GCNs often suffer from high latency and low power efficiency on CPU and GPU platforms due to computational complexity. To address this.....

SeeDetail 

以非局部的解碼器-擠壓-激勵網路及自適應深度列表達成基於編碼器-解碼器的單鏡頭深度估計任務

單鏡頭深度估計是計算機視覺中的一個重要議題。近年來,基於卷積神經網路的編碼器-解碼器架構中展現了合理的結果。在一個強大的編碼器下,人們發現即使是簡單的上採樣過程也能達到良好的準確度....

SeeDetail 

MONOCULAR 3D BASED HUMAN POSE ESTIMATION WITH REFINEMENT BLOCK AND SPECIAL LOSS FUNCTION

In this architecture, we present a 3D HPE by monocular. We use the multi-loss method that depends on 2D heatmaps and volumetric heatmaps and a refinement block to locate root-relative 3D human pose....

SeeDetail 

Multitask Learning on 3D Hand Pose Estimation with Continuous Joints Heatmap

In recent years, deep learning algorithms have been accelerated with GPUs or other volume acceleration hardware, and deep neural networks have gained significant improvements in various tasks....

SeeDetail 

Equipped with Monocular Depth Estimation and Efficient and Accurate Visual Tracking via Cross-Attention Transformer for Human-Following System

With the advancement of technology and the development of human civilization, intelligent robots as "service robots" have been gradually integrated into the practical application of people's daily lives and the good assistants for people’s family life, medical and health, entertainment, and social security. In this paper.....

SeeDetail 

Superpixel 2D to 3D Conversion

此演算法基於superpixel的2D-to-3D,是一個自動深度擷取轉換方法。近來三維立體影像需求的 增加,而三維影像內容資源之缺乏。如果想容易的享受到逼真的立體視覺效果,勢必需要開發低成本、高效率的轉換方 法,將原本二維的影像快速的轉換成三維立體影像。...

SeeDetail  

2D to 3D Conversion

隨著近年來3D電影的風行,如阿凡達(Avatar)與魔境夢遊(Alice in Wonderland)等電影的推出,3D已在全球掀起一股熱潮,然而3D立體顯示技術並非最近幾年才有,早在1928年John Logie Baird就已開發出第一套立體電視系統,而在1935年也已經出現第一部彩色 3D 立體電影,隨後在1950年代美國也拍攝了許多部的3D立體電影。

SeeDetail  

智慧監控平台

本作品是一種前瞻性未來智慧控制系統的模擬應用,相較於市面上的IoT方案大多著點在智慧手機為媒介,我們主要著重在於行為本身的操作,可以避免攜帶式的不便,並且同時擁有更多的未來方案的可行性,在鏡頭模組在居家的重要性逐漸增加的現況中,行為控制與辨識將會是更符合智慧家庭需求的方案。

SeeDetail

Dual-Stream IR-RGB Feature-Fusion Object Detection Network

A dual-stream IR-RGB feature-fusion detection network for robust pedestrian and vehicle detection under challenging lighting. (M.S. thesis, 2026)

基於具有座標注意力和邊緣檢測輔助之雙邊分割網路的實時語義分割任務

語義分割任務在計算機視覺領域中一直是一個重要議題。近年來,卷積神經網路(Convolutional Neural Network)的作法也從比較早期的編碼器-解碼器(Encoder-Decoder)架構,演變至今各種架構都有人使用,對於語義分割任務來說,空間訊息和感受場(receptive field)是不可缺少的,為了使語義分割數方法幾乎都選擇在圖片解析度和低層次的細節訊息上做出妥協,這導致了準確性的大幅下降。在本文中.....

SeeDetail 

A Single-Stage Face Detection and Face Recognition Deep Neural Network Based on Feature Pyramid and Triplet Loss

A practical deep learning face recognition system can be divided into several tasks. These tasks can be time-consuming if executed each task with the original image as the input data....

SeeDetail 

簡易背景更新模型

背景差方法實現運動物體檢測面臨的挑戰主要有:必須適應環境的變化(比如光照的變化造成圖像色度的變化);圖 像中密集出現的物體(比如樹葉或樹幹等密集出現的物體,要正確的檢測出來);必須能夠正確的檢測出背景物體的改 變...

SeeDetail  

低複雜度多物件追蹤系統與閉塞情況的解決方案

多物件偵測與追蹤在電腦視覺領域中是一項重要的研究議題,但絕大多數的物件追蹤演算法的計算過於複雜,並不實 用於即時的追蹤系統中,因此本文提出一個低複雜度的影像多物件追蹤系統,這追蹤主要分成四個部分,第一部分是簡 易背景更新模型的建構,擁有極快速的處理速度,並擁有一定程度的抗噪性,而後續則全部使用前景切割後的黑白圖像 進行處理

SeeDetail  

可即時物件擷取與追蹤之智慧型監控攝影機

在現今的社會裡,人們仰賴監控攝影機來達到許多監控的目的。超級市場需要監控系統、馬路口、或是一般的住家也 都有裝設。然而在一般的監視系統中,有著許多不人性化的地方...

SeeDetail  

Video Surveillance

在現今的社會裡,人們仰賴監控攝影機來達到許多監控的目的。超級市場需要監控系統、馬路口、或是一般的住家也 都有裝設。然而在一般的監視系統中,有著許多不人性化的地方...

SeeDetail  

Video Segmentation

To achieve the content-based functionality, images and video sequences need to be segment as semantic objects by segmentation algorithm. Part of this evolution is due to the need to support a large number of new multimedia applications, such as MPEG-4 and MPEG-7...

SeeDetail  

Intelligent vision on Zigbee

智慧家庭生活是一種趨勢,不僅是年輕一代,在中年老年人間也愈來愈受歡迎。智慧家庭當中,無線網路、感測器和 中央處理器扮演著重要的角色,感測器利用無線網路傳送重要的資訊給中央處理器判斷,也可利用手形辨識,來達成電 器控制,控制環境...

SeeDetail  

應用於高階液晶顯示器之高品質畫面更新率提升轉換技術(FRUC)

為了解決多媒體裝置間大量的連接線路,因此多種家用無線傳輸系統被提出,如wireless HD或者WHDI等。然而現有無線傳輸系統主要專注於點對點高速數據傳輸,而且關於傳輸距離部分也有嚴格限制,因此在使用上會有需多限制。因此一個整合網路機制、無線傳 輸、以及多媒體壓縮之系統被提出,稱之為WHDVIm。在此系統中有兩個重要模組,分別為畫面更新率提昇技術以 及可延伸編解碼。

SeeDetail  

數位家庭之視訊內容即時分析整合系統

在未來的數位家庭中,可經由各種媒介接收多樣化的視訊,為了讓使用者可一邊觀看影片內容,同時也可瞭解其它頻 道的資訊,必須對重要資訊進行擷取整合,由於在影片內容中,常常包含大量的字幕,而這些字幕大多是前後影片內容 的摘要,因此,我們設計一套能自動擷取字幕並能即時進行視訊整合的系統。

SeeDetail  

應用於數位相機之感興趣區域切割嵌入式系統設計(Ti-IEKC6416)

在我們設計的可利用於影像合成之視訊切割系統裡,主要的範圍是應用在數位相機上,在現今的市場中,數位相機可 以說是近年消費性電子產品的大熱門。因此我們應用視訊切割技術,讓數位相機在對人拍攝後可On-line將相片 中之人物擷取出,與感興趣的背景結合、或做更進一步的應用。除此之外、影像是難以在Off-line下自動的去 將ROI (Region of Interest)區域擷取出。我們稱此系統為CameraS( region-of-interest Segmentation for digital Camera)。

SeeDetail  

Video Error Concealment Techniques Using Progressive Interpolation and Boundary Matching Algorithm

*Video compression must be used in this kind of application for bandwidth efficiency.

*The internet network is an error prone environment.

*The packet loss and bit error induces error propagation in the video frame.

SeeDetail