CHAPTER 1 Introduction to Computer Vision and Image Processing Chapter Introduction / Overview Computer vision is the field that enables computers and machines to see, interpret and make decisions from visual data such as images and videos. It combines concepts from digital image processing, mathematics, programming, artificial intelligence and machine learning to extract useful information from pixels. Image processing is the foundation of computer vision. It focuses on improving, transforming and analysing images so that important visual details can be enhanced or measured. Common image processing operations include resizing, cropping, filtering, thresholding, colour conversion and edge detection. OpenCV, the Open Source Computer Vision Library, is one of the most widely used libraries for computer vision. It provides ready-made functions for reading images, processing images, handling videos, detecting objects, extracting features and building real-time vision applications. This chapter introduces the basic building blocks required before students can develop computer vision programs. The chapter begins with the importance and applications of computer vision, then explains how images are represented inside a computer, how pixels and resolution define an image, how colour models work, and how OpenCV can be used to read, display, write, manipulate and inspect images. By the end of this chapter, students will be comfortable with the fundamental vocabulary of computer vision and will be able to write simple OpenCV programs for basic image input, output and manipulation tasks. Learning Outcomes Upon successful completion of this chapter, students will be able to: 1. Define computer vision and image processing and explain the relationship between them. 2. Describe the OpenCV library, its purpose, supported languages and major modules. 3. Identify important features and capabilities of OpenCV for image and video processing. 4. Discuss real-world applications of computer vision in healthcare, industry, agriculture, robotics, surveillance and education. 5. Explain digital image representation using pixels, resolution, bit depth, channels and colour models. 6. Use OpenCV functions to read, display and write images safely with basic error checking. 7. Perform simple image manipulation operations such as cropping, resizing, flipping and region-of-interest extraction. 8. Inspect image properties such as image size, data type, number of channels, shape and total number of pixels. Key Concepts Glossary of Key Terms in This Chapter Computer Vision: A branch of computer science and artificial intelligence that enables machines to acquire, process, analyse and understand visual information from images and videos. Image Processing: The technique of performing operations on an image to enhance it, transform it or extract useful information from it. OpenCV: An open-source computer vision and machine learning library that provides optimized algorithms for image processing, video analysis and real-time vision applications. Pixel: The smallest controllable element of a digital image. Each pixel stores intensity or colour information. Resolution: The number of pixels contained in an image, usually expressed as width x height, such as 1920 x 1080. Colour Model: A mathematical representation used to describe colours, such as RGB, BGR, HSV, Grayscale, YCrCb and Lab. Channel: One component of an image. A grayscale image has one channel, while a colour image usually has three channels. Region of Interest (ROI): A selected portion of an image used for focused processing, such as detecting a face in a cropped area. Image Matrix: A numerical array representation of an image where rows, columns and channels store pixel values. Data Type: The format used to store pixel values, such as uint8 for 8-bit images or float32 for floating-point images. 1.1 Overview of Computer Vision and Image Processing Computer vision and image processing are closely related but not identical. Image processing deals mainly with transforming or improving images, while computer vision goes a step further and attempts to understand what the image contains. For example, increasing the brightness of a low-light photograph is an image processing task, whereas identifying whether the photograph contains a person, vehicle or traffic sign is a computer vision task. In a typical computer vision system, the camera captures an image, the image is cleaned or enhanced, meaningful features are extracted, and a decision is made based on those features. This decision may be simple, such as measuring the size of an object, or complex, such as recognizing a face or guiding a robot. Table 1.1: Computer Vision versus Image Processing Dimension Image Processing Computer Vision Primary Goal Enhance, transform or analyse an image at pixel level. Understand image content and make decisions from visual data. Typical Input Single image or sequence of images. Images, videos, depth maps or multi-camera streams. Output Processed image, enhanced image or measurement. Class label, detected object, coordinates, tracking result or action. Examples Cropping, resizing, denoising, filtering, thresholding. Face recognition, object detection, OCR, autonomous navigation. Main Focus How to improve or transform the image. What the image represents and what action should be taken. Figure 1.1: General computer vision pipeline — capture, preprocess, analyse, interpret and act The success of computer vision depends on the quality of the visual input and the suitability of the algorithm. A clear, well-lit image usually gives better results than a blurred, noisy or poorly exposed image. Therefore, image processing is often used as a preparation step before computer vision tasks. 1.1.1 Basic Stages in a Computer Vision System Main stages of a vision-based solution Image Acquisition: The image or video is captured using a camera, scanner, microscope, drone, satellite or stored file. Preprocessing: The input is resized, cropped, denoised, corrected for brightness or converted to a suitable colour model. Segmentation: The image is divided into meaningful regions, such as foreground and background or object and non-object areas. Feature Extraction: Important characteristics such as edges, corners, texture, colour or shape are identified. Recognition or Measurement: The system classifies objects, detects patterns, measures dimensions or tracks movement. Decision Making: The result is used to trigger an action, generate a report or support human decision-making. 1.1.2 Why Image Processing Is a Foundation for Computer Vision Most vision algorithms work better when the input image is consistent and meaningful. Image processing reduces unwanted variations and highlights useful information. For example, in a medical image, contrast enhancement may make tissues easier to observe. In a traffic camera image, cropping may isolate the road region and remove unnecessary background. Image processing also helps reduce computational cost. A smaller cropped region can be processed faster than a full high-resolution image. Similarly, converting a colour image to grayscale may be sufficient for edge detection and can reduce memory usage. Did You Know? A digital image is not stored as a picture inside a computer. It is stored as numbers arranged in a matrix. Computer vision algorithms work by analysing these numbers. 1.2 Overview of OpenCV and the OpenCV Library OpenCV stands for Open Source Computer Vision Library. It is a large collection of optimized functions and algorithms used for image processing, computer vision and machine learning. OpenCV is popular because it is open source, fast, cross-platform and suitable for both classroom learning and industrial applications. OpenCV was designed to provide common infrastructure for computer vision applications. Instead of writing every image processing algorithm from the beginning, developers can use tested and optimized OpenCV functions. This makes it easier to build projects such as face detection systems, document scanners, traffic monitoring systems and real-time camera applications. Table 1.2: Important Characteristics of OpenCV Characteristic Description Educational Importance Open Source The library can be used, studied and modified under an open-source license. Students can learn from real implementations and build projects freely. Cross-Platform Runs on Windows, Linux, macOS, Android and other environments. The same concepts can be practiced on different systems. Multiple Languages Provides interfaces for C++, Python, Java and other languages. Beginners can use Python while advanced users can use C++ for speed. Optimized Algorithms Includes optimized routines for image processing and vision. Students can focus on concepts before implementing algorithms manually. Real-Time Support Designed for many real-time vision applications. Useful for webcam-based, robotics and surveillance projects. Large Community Extensive documentation, tutorials and community support are available. Troubleshooting and learning become easier. 1.2.1 Installing and Importing OpenCV in Python In Python, OpenCV is usually installed as the opencv-python package and imported using the cv2 module. The common convention is to import it as cv, although many beginner examples also use cv2 directly. # Install OpenCV in a Python environment pip install opencv-python # Import OpenCV in Python import cv2 as cv print(cv.__version__) 1.2.2 Major OpenCV Modules Table 1.3: Major OpenCV Modules and Their Use Module Purpose Common Examples core Basic data structures and mathematical operations. Arrays, matrices, data types, performance utilities. imgcodecs Reading and writing image files. cv.imread(), cv.imwrite(). highgui Displaying images and creating simple windows. cv.imshow(), cv.waitKey(). imgproc Image processing operations. Resize, crop, filter, threshold, edge detection, colour conversion. videoio Reading and writing video streams. Webcam capture, video file input, frame extraction. features2d Feature detection and description. Corners, keypoints, ORB, SIFT-like feature workflows. objdetect Object detection utilities. Cascade classifiers, QR code detection. dnn Deep neural network inference. Using trained models for object detection or classification. Figure 1.2: OpenCV ecosystem — core modules support image input, processing, analysis and output Note For beginners, the most important OpenCV functions in the first unit are cv.imread(), cv.imshow(), cv.imwrite(), image.shape, image.dtype, slicing for cropping, and cv.cvtColor() for colour conversion. 1.3 Features and Capabilities of OpenCV OpenCV is not limited to reading and displaying images. It includes a wide range of functions required for complete vision-based application development. These capabilities range from simple operations such as resizing and drawing shapes to advanced operations such as camera calibration, feature matching, optical flow, object tracking and deep learning model inference. Table 1.4: Features and Capabilities of OpenCV Capability Area Description Example Use Image I/O Read, display and save images in common file formats. Opening a JPEG image and saving the processed result as PNG. Image Transformation Resize, crop, rotate, flip and warp images. Correcting the perspective of a scanned document. Filtering and Enhancement Blur, sharpen, denoise and adjust image contrast. Reducing noise in a low-light camera image. Colour Processing Convert between colour models such as BGR, RGB, HSV and grayscale. Detecting a red object using HSV colour thresholding. Feature Detection Find corners, edges, keypoints and descriptors. Matching the same object in two different images. Video Processing Capture frames from webcams or video files. Building a live attendance or surveillance system. Object Detection Detect faces, objects, shapes or regions of interest. Detecting vehicles in a traffic video. Machine Learning / DNN Use ML algorithms and trained deep learning networks. Running an object detection model on camera frames. Camera Calibration Estimate camera parameters and correct distortion. Improving measurement accuracy in robotics or industrial inspection. 1.3.1 Why OpenCV Is Preferred for Learning OpenCV is preferred in undergraduate courses because it provides a direct connection between theory and practice. Students can immediately observe the effect of operations such as cropping, resizing and colour conversion. The same library can later be extended for advanced topics such as thresholding, edge detection, feature extraction, face detection and deep learning. 1.3.2 Advantages and Limitations of OpenCV Table 1.5: Advantages and Limitations of OpenCV Advantages Limitations / Precautions Fast and optimized for many image processing tasks. Some advanced algorithms require mathematical understanding to use correctly. Large number of functions available in one library. Function names and parameters may confuse beginners initially. Works with Python NumPy arrays, making image data easy to inspect. Images loaded in OpenCV use BGR order by default, not RGB. Supports real-time camera and video processing. Display functions may behave differently in notebooks or server environments. Excellent for prototyping and project development. Deep learning training is usually done in libraries such as TensorFlow or PyTorch, while OpenCV is often used for inference and preprocessing. Best Practice Always read the documentation of a function before using it in a project. Many OpenCV errors are caused by incorrect image path, wrong channel order, unexpected image size or improper data type. 1.4 Applications and Importance of Computer Vision Computer vision is important because visual information is one of the richest sources of data in the real world. Humans naturally understand the world through vision; computer vision attempts to give a similar ability to machines. It is now used in mobile phones, hospitals, factories, vehicles, farms, banks, retail stores and security systems. The importance of computer vision has increased due to the availability of digital cameras, smartphones, drones, cloud computing, GPUs and deep learning models. Many tasks that were once performed manually can now be assisted or automated through computer vision. Table 1.6: Applications of Computer Vision Across Domains Domain Application Importance Healthcare Medical image analysis, tumour detection, X-ray and MRI support. Assists doctors in diagnosis and improves screening efficiency. Automotive Lane detection, traffic sign recognition, pedestrian detection. Supports driver assistance and autonomous vehicle systems. Agriculture Crop disease detection, fruit counting, weed identification. Improves yield monitoring and precision farming. Domain Application Importance Manufacturing Defect detection, product inspection, barcode reading. Improves quality control and reduces manual inspection time. Security Face detection, person tracking, intrusion detection. Enhances monitoring and event detection in public or private spaces. Retail Shelf monitoring, queue analysis, customer movement tracking. Supports inventory management and customer analytics. Education Document scanning, handwriting recognition, virtual labs. Helps automate assessment and digital learning tools. Robotics Object localization, obstacle avoidance, navigation. Allows robots to interact safely with real-world environments. Banking and Documents OCR, cheque processing, identity verification. Automates document-based workflows and reduces errors. 1.4.1 Importance in Modern Society Computer vision improves speed, accuracy and consistency in tasks that involve visual inspection. A trained system can examine thousands of images faster than a human, making it useful in quality control, medical screening and surveillance. However, computer vision should be used responsibly, especially where privacy, fairness and safety are involved. 1.4.2 Importance in Engineering Education For computer science students, computer vision is an excellent field because it combines programming, mathematics, data structures, machine learning and real-world problem solving. Even simple image operations teach important concepts such as arrays, loops, functions, indexing, file input-output and algorithmic thinking. Ethical Reminder Vision systems may process sensitive data such as faces, identities, locations and medical images. Developers must consider privacy, consent, bias, security and the social impact of automated decision-making. 1.5 Image Representation: Pixels, Resolution and Colour Models A digital image is represented as a grid of small units called pixels. Each pixel stores numerical values that describe brightness or colour. In a grayscale image, each pixel usually has one value representing intensity. In a colour image, each pixel usually has three values representing colour channels. 1.5.1 Pixels The word pixel comes from picture element. A pixel is the smallest addressable unit in a digital image. When viewed from a distance, thousands or millions of pixels combine to form a complete image. When zoomed in, individual square-like blocks of colour become visible. In an 8-bit grayscale image, a pixel value usually ranges from 0 to 255. The value 0 represents black, 255 represents white, and values between them represent different shades of gray. In a colour image, each channel may have a separate value from 0 to 255. Table 1.7: Pixel Values in Common Image Types Image Type Channels Typical Pixel Value Meaning Binary image 1 0 or 1 / 0 or 255 Black and white image used for masks and segmentation. Grayscale image 1 0 to 255 Single intensity value per pixel. BGR / RGB colour image 3 (Blue, Green, Red) or (Red, Green, Blue) Three colour components per pixel. RGBA / BGRA image 4 Three colour channels + alpha Colour image with transparency information. Floating-point image 1 or more Often 0.0 to 1.0 or other numeric range Used in scientific imaging and deep learning pipelines. 1.5.2 Image Resolution Image resolution refers to the number of pixels present in an image. It is usually written as width x height. For example, an image with resolution 1920 x 1080 has 1920 pixels in each row and 1080 pixels in each column direction. The total number of pixels is 1920 x 1080 = 2,073,600 pixels. Higher resolution images contain more detail but require more storage, memory and computation. For many computer vision tasks, images are resized to a standard dimension so that algorithms receive consistent input. Table 1.8: Resolution, Total Pixels and Typical Use Resolution Total Pixels Common Name / Use 640 x 480 0.31 million VGA; simple webcam experiments and low-cost processing. 1280 x 720 0.92 million HD; video streaming and basic surveillance. 1920 x 1080 2.07 million Full HD; common video and camera resolution. 3840 x 2160 8.29 million 4K; high-detail video and professional imaging. 224 x 224 50,176 Common deep learning input size for many image classification models. 1.5.3 Image Matrix Representation In Python OpenCV, an image is represented as a NumPy array. For a grayscale image, the array is two-dimensional with shape (height, width). For a colour image, the array is three-dimensional with shape (height, width, channels). The first dimension represents rows, the second dimension represents columns, and the third dimension represents colour channels. # Shape of a colour image loaded with OpenCV # height = number of rows, width = number of columns, channels = 3 image.shape # Example output: (480, 640, 3) Figure 1.3: Matrix representation of a colour image — height, width and channels 1.5.4 Colour Models A colour model defines how colours are represented numerically. Different colour models are useful for different tasks. OpenCV loads most colour images in BGR format by default, whereas many plotting libraries and image standards commonly use RGB ordering. This difference is a common source of confusion for beginners. Table 1.9: Common Colour Models in Image Processing Colour Model Components Typical Use Grayscale Intensity only Edge detection, thresholding, document processing and simple measurements. RGB Red, Green, Blue Standard display model used in monitors, cameras and many image libraries. BGR Blue, Green, Red Default channel order used by OpenCV for colour images. HSV Hue, Saturation, Value Colour segmentation and object tracking based on colour. YCrCb / YCbCr Luminance and chrominance components Video compression, skin colour detection and image coding. Lab Lightness and colour-opponent dimensions Colour correction and perceptual colour analysis. CMYK Cyan, Magenta, Yellow, Black Printing and publishing workflows. 1.5.5 BGR versus RGB When an image is read using OpenCV, the colour channels are usually arranged in BGR order. If the same image is displayed using a library that expects RGB, the colours may appear incorrect. Therefore, BGR to RGB conversion is often required before displaying OpenCV images with Matplotlib. import cv2 as cv import matplotlib.pyplot as plt img_bgr = cv.imread('flower.jpg') img_rgb = cv.cvtColor(img_bgr, cv.COLOR_BGR2RGB) plt.imshow(img_rgb) plt.axis('off') plt.show() Important Point In OpenCV, the image shape is usually written as (height, width, channels), but image resolution is commonly written as width x height. This difference should be remembered during programming. 1.6 Reading, Writing and Displaying Images Reading, displaying and writing images are the first practical operations in OpenCV. These operations allow a program to load an image from disk, show it to the user and save the processed output. Although the functions are simple, correct handling of paths, file formats and error checking is essential. 1.6.1 Reading Images with cv.imread() The function cv.imread() loads an image from a specified file path. If the path is wrong or the file cannot be read, the function returns None in Python. Therefore, every beginner program should check whether the image has been loaded successfully before processing it. import cv2 as cv img = cv.imread('sample.jpg') if img is None: print('Error: Image could not be loaded. Check the file path.') else: print('Image loaded successfully') print(img.shape) Table 1.10: Common Image Reading Flags in OpenCV Flag Meaning Use Case cv.IMREAD_COLOR Loads a colour image in BGR format. Alpha channel is ignored. Default choice for most colour image programs. cv.IMREAD_GRAYSCALE Loads image as a single-channel grayscale image. Useful for edge detection, thresholding and document analysis. Flag Meaning Use Case cv.IMREAD_UNCHANGED Loads image as it is, including alpha channel if available. Useful for PNG images with transparency. 1.6.2 Displaying Images with cv.imshow() The function cv.imshow() displays an image in a window. The function cv.waitKey() is used to keep the window open until a key is pressed or for a specified number of milliseconds. The function cv.destroyAllWindows() closes OpenCV display windows. import cv2 as cv img = cv.imread('sample.jpg') if img is not None: cv.imshow('Display Window', img) cv.waitKey(0) # waits until a key is pressed cv.destroyAllWindows() # closes all OpenCV windows 1.6.3 Writing Images with cv.imwrite() The function cv.imwrite() saves an image to a file. The file extension decides the output format, such as .jpg, .png or .bmp. It is good practice to check the return value of cv.imwrite(), because it returns True if the image was saved successfully and False otherwise. import cv2 as cv img = cv.imread('sample.jpg') if img is not None: success = cv.imwrite('output.png', img) if success: print('Image saved successfully') else: print('Image could not be saved') Table 1.11: Image I/O Functions in OpenCV Function Purpose Important Notes cv.imread(path, flag) Reads an image from disk. Returns None if the file cannot be read. cv.imshow(window_name, image) Displays image in a named window. Usually used with cv.waitKey(). cv.waitKey(delay) Waits for a keyboard event. 0 means wait indefinitely. cv.destroyAllWindows() Closes all display windows. Used after displaying images. Function Purpose Important Notes cv.imwrite(path, image) Saves image to disk. Returns True if saved. 1.6.4 Common Errors While Reading Images Table 1.12: Common Image Loading Errors and Solutions Problem Likely Cause Solution img is None Wrong file path, missing file or unsupported format. Check spelling, folder location and file extension. Colours look different BGR image displayed as RGB. Convert BGR to RGB before using Matplotlib. Window closes immediately cv.waitKey() not used. Add cv.waitKey(0) after cv.imshow(). Image not saved Invalid output path or no write permission. Use a valid folder and check cv.imwrite() return value. Program works in Python IDE but not in notebook GUI display limitations. Use Matplotlib or notebook display functions instead of cv.imshow(). Laboratory Tip Keep the image file in the same folder as the Python program during beginner experiments. This reduces path errors and makes programs easier to test. 1.7 Image Manipulation and Cropping Image manipulation means changing an image by performing operations such as cropping, resizing, rotating, flipping, changing brightness, drawing shapes or extracting a region of interest. In OpenCV Python, many simple manipulations can be done using NumPy array indexing because an image is stored as an array. 1.7.1 Cropping an Image Cropping means selecting a rectangular portion of an image. The cropped portion is also called a region of interest. In Python, cropping can be performed using array slicing. The row range represents the y-direction and the column range represents the x-direction. import cv2 as cv img = cv.imread('sample.jpg') # Crop using row and column ranges # Format: image[y1:y2, x1:x2] crop = img[100:300, 150:400] cv.imshow('Cropped Region', crop) cv.waitKey(0) cv.destroyAllWindows() Remember In image slicing, the order is image[y1:y2, x1:x2], not image[x1:x2, y1:y2]. Rows correspond to the y-axis and columns correspond to the x-axis. 1.7.2 Resizing Images Resizing changes the width and height of an image. It is commonly used to reduce memory usage, standardize image dimensions, prepare input for machine learning models or display large images conveniently. import cv2 as cv img = cv.imread('sample.jpg') # Resize to a fixed width and height resized = cv.resize(img, (300, 200)) # (width, height) # Resize by scale factors smaller = cv.resize(img, None, fx=0.5, fy=0.5) 1.7.3 Flipping and Rotating Images Flipping creates a mirror image horizontally, vertically or in both directions. Rotation changes the orientation of the image. These operations are useful in data augmentation, camera correction and visual effects. import cv2 as cv img = cv.imread('sample.jpg') horizontal_flip = cv.flip(img, 1) # 1 = horizontal vertical_flip = cv.flip(img, 0) # 0 = vertical both_flip = cv.flip(img, -1) # -1 = both directions rotated = cv.rotate(img, cv.ROTATE_90_CLOCKWISE) 1.7.4 Drawing on Images OpenCV provides functions to draw lines, rectangles, circles and text on images. These functions are useful for marking detected objects, visualizing regions of interest and creating annotated outputs. import cv2 as cv img = cv.imread('sample.jpg') # Draw a rectangle: (image, top-left, bottom-right, colour, thickness) cv.rectangle(img, (50, 50), (250, 200), (0, 255, 0), 2) # Put text on image cv.putText(img, 'Object', (50, 40), cv.FONT_HERSHEY_SIMPLEX, 1, (0, 255, 0), 2) Table 1.13: Basic Image Manipulation Operations Operation OpenCV / Python Method Purpose Crop image[y1:y2, x1:x2] Extract a selected region of interest. Resize cv.resize(image, size) Change image dimensions. Flip cv.flip(image, flipCode) Mirror image horizontally or vertically. Rotate cv.rotate(image, rotateCode) Rotate image by fixed angles. Colour conversion cv.cvtColor(image, code) Convert between BGR, RGB, grayscale, HSV and other models. Drawing cv.rectangle(), cv.circle(), cv.line(), cv.putText() Annotate image or mark detections. Figure 1.4: Region of interest extraction using row and column slicing 1.8 Understanding Image Properties: Size, Type and Channels Before processing an image, it is important to inspect its properties. Image properties tell us the dimensions, number of channels, number of pixels and data type. These details determine how much memory the image uses and what operations can be safely applied. 1.8.1 Shape, Size and Data Type In OpenCV Python, image.shape gives the dimensions of the image. For colour images it usually returns (height, width, channels). image.size gives the total number of values in the array, and image.dtype gives the data type used to store pixel values. import cv2 as cv img = cv.imread('sample.jpg') if img is not None: print('Shape:', img.shape) # (height, width, channels) print('Size:', img.size) # total number of values print('Data type:', img.dtype) # usually uint8 Table 1.14: Image Properties in OpenCV Property Example Output Meaning image.shape (480, 640, 3) Height = 480, width = 640, channels = 3. image.size 921600 Total number of values: 480 x 640 x 3. image.dtype uint8 Each pixel/channel value is stored as an 8-bit unsigned integer. image.ndim 3 Number of dimensions in the array. image[0, 0] [B, G, R] Pixel value at first row and first column. 1.8.2 Understanding Channels A channel is a layer of information in an image. A grayscale image has one channel. A BGR image has three channels: blue, green and red. A transparent image may have four channels, where the fourth channel is the alpha channel. Understanding channels is important because many operations require a specific number of channels. Table 1.15: Channel Interpretation Shape Image Type Interpretation (height, width) Grayscale One intensity value per pixel. (height, width, 3) BGR / RGB colour Three colour values per pixel. (height, width, 4) BGRA / RGBA colour Three colour values plus transparency. (number_of_frames, height, width, channels) Video / image batch Multiple images stored together. 1.8.3 Memory Requirement of an Image The memory required by an image depends on width, height, number of channels and bytes per channel. An 8-bit colour image of resolution 1920 x 1080 has 1920 x 1080 x 3 values. Since each uint8 value uses one byte, the approximate memory is about 6.22 MB before compression. height, width, channels = img.shape bytes_used = img.size * img.itemsize print('Approximate memory in bytes:', bytes_used) 1.8.4 Summary of Beginner Rules Beginner Rules for OpenCV Image Handling 1. Always check whether cv.imread() returned a valid image before processing. 2. Remember that OpenCV stores colour images in BGR order by default. 3. Use image.shape to understand height, width and channels. 4. Use image[y1:y2, x1:x2] for cropping, where y represents rows and x represents columns. 5. Resize large images when computation speed or memory is a concern. 6. Convert colour models carefully before using functions that expect grayscale, RGB or HSV input. 7. Save outputs with cv.imwrite() and check whether the operation succeeded. 1.9 Further Reading / Viewing The following resources are recommended for students who want to strengthen their understanding of OpenCV, computer vision and image processing fundamentals. Table 1.16: Recommended Resources Resource Type Resource Reason for Recommendation Official Documentation OpenCV Documentation and Tutorials Primary reference for functions, parameters and examples. Official Website OpenCV.org About and Library Pages Useful overview of OpenCV capabilities, licensing and applications. Textbook Digital Image Processing by Gonzalez and Woods Classic textbook for image processing theory. Textbook Learning OpenCV by Bradski and Kaehler Practical reference for OpenCV concepts and applications. Online Practice OpenCV Python Tutorials Hands-on examples for reading images, colour conversion and basic operations. Video Learning Computer Vision and OpenCV beginner tutorials Helpful for visual demonstrations of image manipulation tasks. Suggested Viewing Students may practice this unit by selecting any five images, loading them in OpenCV, printing their properties, cropping a region of interest, converting them to grayscale and saving the output files. 1.10 Assessment Questions Part A - Multiple Choice Questions (1 Mark Each) Q1. Which library is commonly used for computer vision and image processing in Python? 1. NumPy only 2. OpenCV 3. MySQL 4. HTML Answer: (b) Q2. What is the smallest unit of a digital image called? 1. Frame 2. Pixel 3. Channel 4. Kernel Answer: (b) Q3. Which OpenCV function is used to read an image from a file? 1. cv.read() 2. cv.imread() 3. cv.loadImage() 4. cv.open() Answer: (b) Q4. What does cv.imshow() do? 1. Saves an image 2. Displays an image in a window 3. Crops an image 4. Changes image colour Answer: (b) Q5. Which image property returns height, width and channels in Python OpenCV? 1. image.type 2. image.shape 3. image.length 4. image.path Answer: (b) Q6. What is the default colour channel order used by OpenCV for colour images? 1. RGB 2. CMYK 3. BGR 4. HSV Answer: (c) Q7. Which slicing format is correct for cropping an OpenCV image? 1. image[x1:x2, y1:y2] 2. image[y1:y2, x1:x2] 3. image[channel, row, col] 4. image(width, height) Answer: (b) Q8. What does a grayscale image usually contain? 1. Three colour channels 2. One intensity channel 3. Four alpha channels 4. No pixel values Answer: (b) Part B - Short Answer Questions (5 Marks Each) Q9. Define computer vision and image processing. Explain the difference between them with one example each. Q10. What is OpenCV? List any five important features of OpenCV. Q11. Explain the terms pixel, resolution, channel and colour model. Q12. Differentiate between RGB, BGR, grayscale and HSV colour models. Q13. Write a Python OpenCV program to read an image, display it and save it with a new filename. Q14. Explain why it is important to check whether cv.imread() has loaded an image successfully. Q15. What is image cropping? Explain the format image[y1:y2, x1:x2]. Q16. Explain image.shape, image.size and image.dtype with examples. Part C - Long Answer / Essay Questions (10 Marks Each) Q17. Explain the complete workflow of a basic computer vision system from image acquisition to decision making. Support your answer with a suitable example such as face detection or traffic sign recognition. Q18. Discuss the major applications of computer vision in healthcare, agriculture, manufacturing, autonomous vehicles and security. Explain the importance of computer vision in modern society. Q19. Explain digital image representation in detail. Your answer should include pixels, resolution, bit depth, image matrix, colour channels, grayscale images and colour models. Q20. Describe the basic OpenCV image input-output workflow. Include cv.imread(), cv.imshow(), cv.waitKey(), cv.destroyAllWindows() and cv.imwrite() with suitable code examples. Q21. Explain image manipulation operations in OpenCV, including cropping, resizing, flipping, rotation, drawing and colour conversion. Give syntax and examples wherever necessary. Part D - Analytical / Case-Based Questions Q22 (Case Study). A college wants to build a simple system that captures student ID card images and automatically crops the face region for further processing. Explain which OpenCV operations will be needed for reading the image, checking image properties, cropping the region of interest, resizing the crop and saving the output.