Court-Shot detection in tennis broadcast video
Ahmed Shihab, Krzysztof Zienowicz · Research Repository (Kingston University London) · 2007
We propose a simple method of filtering sports video data so that play frames are retained. We built a model of similar camera frames using a Gaussian Mixture Model and then classified new frames on the basis of that model; the results were very accurate. Our method uses a two dimensional Gaussian weighting function for better and faster separation between play and non-play shots. We used the UV colour subspace of the YUV colour space, which showed insensitivity to brightness changes and allowed realtime processing. The experiments demonstrate the accuracy of our proposed method for shot detection and classification as tested on broadcast TV tennis footage. 1 Shot-Type Recognition Sports broadcasts contain a type of camera shot1 that is just high enough, and just wideangled enough, to capture the events on the field clearly. We use the term play to distinguish this type of camera shot from non-play shots like audience scenes, player zooms, or commercials. Play shots are focussed on the action in the field and are suitable for locating players with respect to the field boundaries. Note that they do not necessarily come from the same camera; several cameras may provide this type of shot. Analysis of sports video footage depends on locating the players with respect to the field borders; camera shots that zoom on players need to be removed as should shots that focus on the audience, or other non-play events. In this paper, we present a method to remove non-play shots from Tennis match recordings; this is a necessary preliminary step to subsequent tracking work; we present an effective solution and explore its limits. A summary of our solution is as follows: we estimated a Gaussian Mixture Model (GMM) for the ”play” shots from training sequences. During training, we sampled pixel locations from a Gaussian with mean equal to the image center and standard deviation equal to 1/6th the corresponding dimensions. Given an input image, the probability that each pixel from the image came from the ”play” GMM is weighted by its distance from the image center and summed. Based on this score, the frame is classified as ”play” or ”non-play”. Experimental results are presented on TV broadcast video. 1A shot is a video sequence that consists of video frames from one camera captured continuously, i.e., at the camera’s normal capture rate. Figure 1: From the left: a) training frame, b) colour data in YUV colour space, c) colour data projected on UV plane, d) GMM for considered frame on UV plane (each ellipse represents distance equal one eigenvalue, line width represents prior value). Our method does not simply verify if a pixel is green or not, then sum over all the pixels. The GMM is necessary in order to generalise the method to any tennis or sports imagery. This problem is fundamentally similar to the content-based retrieval (CBR) problem. In CBR, the goal is to find the images, in a database, that are most similar to the one in hand (this is the query by example case). Analogously, our goal is to find the frames, in a video sequence, that are most similar to our set of play frames (the training set). Whereas our problem is easier than the general CBR problem because we only look for one class of images, the play class; everything else is labelled non-play. Furthermore, our domain is very restricted; our ”database” contains video footage from only one sports event.