[{"content":"A quick note on Graph RAG and when it beats plain vector search.\nThe problem with plain RAG Standard RAG embeds document chunks into vectors and retrieves the chunks most similar to a question. It works well when the answer lives inside one or two chunks. It fails on multi-hop questions — questions where the answer is spread across several documents and requires following relationships. Example: \u0026ldquo;Which suppliers ship parts that failed quality checks in plants served by our Hamburg warehouse?\u0026rdquo; No single chunk answers that.\nWhat Graph RAG adds Graph RAG builds a knowledge graph over the corpus: entities (suppliers, parts, warehouses) become nodes, relationships (ships-to, supplies, failed) become edges. Retrieval then combines:\nGraph traversal — walk the relationships to collect connected facts Vector search — retrieve semantically similar chunks as usual The LLM answers from both the raw text and the structured neighborhood around the relevant entities.\nTypical pipeline Indexing: an LLM extracts entities and relationships from each document → graph store (e.g. Neo4j, Kùzu, or an in-memory NetworkX graph for experiments) Query time: extract entities from the question → find them in the graph → expand a neighborhood (1–2 hops) → feed subgraph + retrieved chunks into the prompt When it\u0026rsquo;s worth it Yes: questions that join facts across documents, \u0026ldquo;how are A and B related\u0026rdquo; queries, global summaries (\u0026ldquo;what themes run through all these reports?\u0026rdquo; — this is what Microsoft\u0026rsquo;s GraphRAG is optimized for) No: simple lookups, small corpora, when a plain vector index already answers well — graph extraction adds indexing cost and a second system to maintain Rule of thumb Start with vector RAG. Add the graph layer when eval questions show answers requiring chained relationships the retriever keeps missing.\n","permalink":"https://naeem-bebit.github.io/posts/graph-rag/","summary":"How knowledge graphs improve retrieval-augmented generation","title":"Graph RAG: Adding Structure to Retrieval"},{"content":"In this post I would like to explain on what is Docker and why we need Docker\nChapter 1: Docker for Data Science In data science, where reproducibility and dependency management are crucial, Docker emerges as a powerful solution. Docker is an open-source platform that automates application deployment in portable containers. It ensures consistent and hassle-free execution across various environments.\nWhy Docker Matters Reproducibility: Docker replicates your data science environment precisely, eliminating the \u0026ldquo;it works on my machine\u0026rdquo; issue.\nDependency Management: Encapsulate project dependencies in a container, avoiding conflicts and enabling multiple projects to coexist.\nPortability: Docker containers run seamlessly on any platform, simplifying collaboration and deployment.\nEfficiency: Lightweight containers share system resources efficiently, enabling multiple concurrent workloads.\nIsolation: Each Docker container operates independently, accommodating diverse libraries and operating systems.\nAdvantages of Docker in Data Science Rapid Setup: Docker containers spin up in seconds, reducing setup time.\nCollaboration: Easily share Docker images for reproducible experiments and collaborative work.\nScalability: Docker integrates seamlessly with cluster computing environments like Kubernetes.\nVersion Control: Docker images can be versioned, tracking changes to your data science environment.\nIn the upcoming chapters, we\u0026rsquo;ll explore Docker\u0026rsquo;s practical use in data science, including custom image creation, efficient container management, and real-world applications. Docker is more than a tool; it\u0026rsquo;s a transformative approach to data science development and deployment. Welcome to the Docker revolution!\nSo, fasten your seatbelts and get ready to embark on a journey that will transform the way you work with data in the world of data science. Welcome to the Docker revolution!\nChapter 2: Getting Started with Docker Now that you understand why Docker is a game-changer in the world of data science, let\u0026rsquo;s dive into the practical side of things. In this chapter, we\u0026rsquo;ll walk you through the essential steps to get started with Docker, from installation to running your first container.\nInstalling Docker Before you can start working with Docker, you need to install it on your system. Fortunately, Docker provides easy-to-follow installation instructions for various operating systems:\nDocker Desktop for Windows Docker Desktop for macOS Docker for Linux Follow the instructions for your specific operating system to install Docker. Once installed, you\u0026rsquo;ll have access to the Docker command-line interface (CLI) and Docker Dashboard (if you\u0026rsquo;re using Docker Desktop).\nRunning Your First Docker Container With Docker installed, you\u0026rsquo;re ready to create and run your first Docker container. Here\u0026rsquo;s a simple example to get you started:\nPull an official Docker image (e.g., the \u0026ldquo;hello-world\u0026rdquo; image) docker pull hello-world Run a container from the image docker run hello-world This basic example demonstrates how easy it is to pull an image from Docker Hub (the default image repository) and run a container. The \u0026ldquo;hello-world\u0026rdquo; container will print a friendly message to your terminal to confirm that Docker is working correctly.\nUnderstanding Docker Images and Containers Before we delve deeper into Docker, it\u0026rsquo;s crucial to understand two fundamental concepts: Docker images and containers.\nDocker Image: An image is a lightweight, standalone, and executable package that includes everything needed to run a piece of software, including the code, runtime, libraries, and system tools.\nDocker Container: A container is a running instance of a Docker image. It\u0026rsquo;s an isolated environment that runs the software contained in the image.\nIn the upcoming chapters, you\u0026rsquo;ll learn how to create custom Docker images tailored to your data science projects, manage containers efficiently, and use Docker to tackle real-world data science challenges.\nNow that you have Docker installed and have run your first container, you\u0026rsquo;re ready to explore the practical applications of Docker in the data science field. Let\u0026rsquo;s continue our journey into the world of Docker!\nChapter 3: Docker Basics for Data Scientists In this chapter, we will delve deeper into the fundamental aspects of Docker that are essential for data scientists. Understanding these core concepts will enable you to harness the full power of Docker in your data science projects. We\u0026rsquo;ll cover how to create custom Docker images tailored to your specific data science needs, efficiently manage containers, and leverage Docker to enhance your data science workflow.\nDockerfile: Building Custom Docker Images A Dockerfile is a script that contains a set of instructions for creating a custom Docker image. As a data scientist, Dockerfiles empower you to encapsulate your entire data science environment, including dependencies, libraries, and even your code, into a portable container. In this section, we\u0026rsquo;ll explore:\nWriting Dockerfiles: We\u0026rsquo;ll guide you through the process of creating a Dockerfile for your data science project. You\u0026rsquo;ll learn how to specify the base image, install necessary packages, and set up your working environment.\nBuilding Custom Images: Once you have your Dockerfile ready, we\u0026rsquo;ll demonstrate how to use it to build custom Docker images. This process allows you to capture the exact state of your data science environment, making it highly reproducible.\nBest Practices: To ensure the efficiency and reproducibility of your Docker images, we\u0026rsquo;ll provide you with best practices and optimization tips for writing Dockerfiles tailored to data science workflows.\nDocker Compose: Managing Multi-Container Applications Data science projects often involve multiple interconnected services and containers that need to work together seamlessly. Docker Compose is a powerful tool that allows you to define, configure, and run multi-container Docker applications with ease. In this section, we will cover:\nCreating Docker Compose Files: We\u0026rsquo;ll guide you through the process of creating Docker Compose files, which define your application\u0026rsquo;s services and their configurations. You\u0026rsquo;ll learn how to specify dependencies, set environment variables, and establish network connections.\nRunning Multi-Container Applications: You\u0026rsquo;ll discover how to use Docker Compose to start and manage your services as a unified application stack. This makes orchestrating complex data science workflows much more manageable.\nOrchestrating Data Science Workflows: We\u0026rsquo;ll provide practical examples of how Docker Compose can streamline your data science workflow. From setting up data ingestion pipelines to deploying machine learning models, Docker Compose can simplify the coordination of various components.\nDocker Volumes: Managing Data Persistence Data is at the core of data science, and effectively managing data within Docker containers is paramount. Docker volumes offer a solution for persisting data outside of containers, ensuring that your valuable datasets, model outputs, and other critical information are retained. In this section, we\u0026rsquo;ll explore:\nUnderstanding Docker Volumes: We\u0026rsquo;ll delve into how Docker volumes work and why they are essential in data science. You\u0026rsquo;ll gain insights into data persistence mechanisms within containers.\nUsing Volumes for Data Persistence: Practical guidance on creating and managing Docker volumes for your containers. This includes strategies for handling data in scenarios where data durability and persistence are crucial.\nBest Practices for Data Management: We\u0026rsquo;ll discuss best practices and data management strategies specific to Dockerized data science projects. You\u0026rsquo;ll learn how to structure your data storage to balance performance, scalability, and data integrity.\nBy the end of this chapter, you\u0026rsquo;ll possess a strong foundation in Docker\u0026rsquo;s core concepts and understand how to apply them effectively to your data science work. You\u0026rsquo;ll be well-equipped to build custom Docker images tailored to your projects, efficiently manage multi-container applications with Docker Compose, and ensure data persistence using Docker volumes.\nWith these skills, you\u0026rsquo;ll be ready to unlock the full potential of Docker in your data science endeavors. Let\u0026rsquo;s embark on this journey and explore the Docker basics that are essential for data scientists.\nChapter 4: Containerizing Data Science Environments In this chapter, we\u0026rsquo;ll take a deeper dive into the practical application of Docker in the data science field. Specifically, we\u0026rsquo;ll explore how to containerize your data science environments using Docker. Containerization is a powerful technique that allows you to encapsulate your entire data science setup, including libraries, dependencies, and code, within a Docker container. This approach not only enhances reproducibility but also simplifies collaboration and deployment in data science projects.\nCreating Docker Images for Data Science Tools Data scientists rely on a wide range of tools and libraries to perform tasks such as data preprocessing, analysis, machine learning, and visualization. Docker enables you to create customized images containing these tools, ensuring consistency across your team and environments. In this section, we\u0026rsquo;ll cover:\nSelecting Base Images: How to choose the right base image for your data science needs.\nInstalling Libraries: Techniques for installing data science libraries and dependencies within your Docker image.\nIncluding Code: Strategies for adding your data science project code to the container image.\nOptimizing Image Size: Best practices for minimizing the size of your Docker images to improve efficiency and speed.\nBuilding Versatile Data Science Environments One of the great advantages of Docker is its ability to create isolated and versatile environments. This section will delve into creating Docker images that serve different data science purposes. We\u0026rsquo;ll explore:\nPython Environments: How to build Docker images tailored for Python-based data science projects.\nR Environments: Creating Docker containers for R users, including the installation of R libraries and packages.\nJupyter Notebooks: Leveraging Docker to set up Jupyter Notebook environments for interactive data exploration and analysis.\nMachine Learning Frameworks: Building Docker images that come pre-configured with popular machine learning frameworks like TensorFlow and PyTorch.\nBest Practices for Data Science Docker Images Ensuring that your Docker images are well-optimized, maintainable, and secure is essential for seamless data science workflows. In this section, we\u0026rsquo;ll discuss:\nVersion Control: Strategies for maintaining version control of your Docker images.\nReproducibility: How to create reproducible Docker images for your data science projects.\nSecurity Considerations: Best practices for securing your Dockerized data science environments.\nDocumentation: The importance of documenting your Docker images for your team and future reference.\nBy the end of this chapter, you\u0026rsquo;ll have a deep understanding of how to use Docker to containerize your data science environments effectively. You\u0026rsquo;ll be able to create custom Docker images tailored to your data science tools, libraries, and projects, making it easier than ever to maintain consistency, collaborate with colleagues, and deploy data science solutions. Containerization is a game-changer in data science, and you\u0026rsquo;re on your way to mastering it.\nChapter 5: Using Docker in Data Science Projects In the previous chapters, we\u0026rsquo;ve explored the fundamentals of Docker and how to containerize your data science environments. Now, it\u0026rsquo;s time to put that knowledge into action. In this chapter, we\u0026rsquo;ll delve into the practical aspects of using Docker in your data science projects.\nSetting Up a Data Science Project with Docker Starting a data science project with Docker involves more than just creating a container. It\u0026rsquo;s about structuring your project for efficiency, reproducibility, and collaboration. In this section, we\u0026rsquo;ll guide you through:\nProject Organization: Best practices for structuring your data science project directory to work seamlessly with Docker.\nDocker Compose for Projects: How to define and manage multi-container setups for your data science projects using Docker Compose.\nEnvironment Variables: Leveraging environment variables within your containers to adapt your environment dynamically.\nVersion Controlling Docker Configurations In the world of data science, version control is essential not only for your code and data but also for your Docker configurations. In this section, we\u0026rsquo;ll cover:\nGit Integration: How to integrate your Docker-related files and configurations into your Git version control workflow.\nDocker Image Versioning: Strategies for versioning your Docker images to ensure reproducibility.\nContinuous Integration (CI) with Docker: Incorporating Docker into CI/CD pipelines to automate testing and deployment of your data science projects.\nCollaborating with Others Using Docker Collaboration is at the heart of many data science projects, and Docker can facilitate seamless teamwork. In this section, we\u0026rsquo;ll explore:\nSharing Docker Images: Strategies for sharing Docker images with team members and collaborators.\nCollaborative Environments: Setting up collaborative environments using Docker Compose for joint experimentation and development.\nReproducibility in Collaboration: Ensuring that your collaborators can easily reproduce your work by providing Docker-based project environments.\nTroubleshooting and Debugging Even with the best preparations, issues can arise. This section will help you handle them effectively:\nDocker Logs: How to access and interpret container logs for debugging.\nTroubleshooting Common Problems: Tips and common solutions for resolving Docker-related issues in data science projects.\nPerformance Optimization: Techniques for optimizing the performance of your Dockerized data science projects.\nBy the end of this chapter, you\u0026rsquo;ll be well-equipped to integrate Docker seamlessly into your data science projects. Whether you\u0026rsquo;re working on individual experiments or collaborating with a team, Docker will empower you to create reproducible, efficient, and collaborative data science environments. Let\u0026rsquo;s dive into the practical applications of Docker in your data science work!\nChapter 6: Orchestrating Docker in Data Pipelines In the world of data science, data pipelines are the backbone of many projects. These pipelines handle data ingestion, processing, transformation, and model training, making them a critical part of the data science workflow. In this chapter, we\u0026rsquo;ll explore how Docker can be effectively used to orchestrate data pipelines, ensuring scalability, reproducibility, and ease of management.\nDocker Swarm and Kubernetes for Scaling Workloads Data pipelines often involve a multitude of tasks and processes that need to be orchestrated seamlessly. Docker Swarm and Kubernetes are two popular container orchestration tools that can help you scale and manage your data science workloads efficiently. We\u0026rsquo;ll cover:\nDocker Swarm: How to use Docker Swarm to manage a cluster of Docker hosts and deploy services for your data pipelines.\nKubernetes: An introduction to Kubernetes for container orchestration, including the deployment of data science workloads.\nScaling Data Pipelines: Strategies for scaling data pipelines using container orchestration to handle large datasets and computationally intensive tasks.\nManaging Distributed Data Processing with Docker Data science often involves distributed data processing, where data is processed across multiple nodes or containers to handle large-scale datasets. Docker can simplify the management of distributed systems in data pipelines. We\u0026rsquo;ll explore:\nDistributed Computing with Docker: Techniques for setting up distributed data processing using Docker containers.\nBig Data Tools: How to integrate Docker with big data tools such as Apache Spark and Hadoop for distributed data processing.\nParallel Processing: Leveraging Docker to parallelize data processing tasks and optimize performance.\nReal-World Examples of Docker in Data Pipelines To illustrate the practical application of Docker in data pipelines, we\u0026rsquo;ll dive into real-world use cases. You\u0026rsquo;ll discover how organizations use Docker to streamline their data processing workflows, including:\nData Ingestion and ETL: Using Docker to create flexible and scalable data ingestion pipelines.\nBatch Processing: Implementing batch processing pipelines with Docker containers for large-scale data processing.\nReal-time Data Streams: Orchestrating Docker containers to process real-time data streams and generate insights in real-time.\nBy the end of this chapter, you\u0026rsquo;ll have a comprehensive understanding of how Docker can be leveraged to orchestrate data pipelines, whether you\u0026rsquo;re working with distributed data processing, big data tools, or real-time data streams. You\u0026rsquo;ll be equipped to design and manage efficient and scalable data pipelines for your data science projects, taking full advantage of containerization technologies.\nChapter 7: Docker in Production for Data Science Up to this point, we\u0026rsquo;ve explored how Docker can benefit data science development and experimentation. Now, it\u0026rsquo;s time to transition from development to production. In this chapter, we\u0026rsquo;ll dive into the best practices and considerations for using Docker in a production environment for data science applications.\nProductionizing Data Science Workflows Productionizing data science workflows requires a different set of considerations compared to development and experimentation. In this section, we\u0026rsquo;ll cover:\nInfrastructure as Code (IaC): Leveraging tools like Terraform and Ansible to automate infrastructure provisioning and management.\nHigh Availability: Designing and deploying highly available Docker-based data science applications to ensure reliability and uptime.\nScalability: Strategies for scaling Dockerized data science applications to handle increased workloads.\nMonitoring and Logging Effective monitoring and logging are essential for maintaining and troubleshooting data science applications in production. We\u0026rsquo;ll explore:\nMonitoring Tools: An overview of monitoring tools like Prometheus and Grafana for tracking container performance and health.\nLogging Best Practices: Implementing logging best practices to capture and analyze application logs.\nAlerting: Setting up alerting systems to proactively address issues in your production environment.\nSecurity Considerations Security is a top priority when deploying data science applications in production. We\u0026rsquo;ll discuss:\nContainer Security: Best practices for securing Docker containers, including image scanning and vulnerability assessment.\nAccess Control: Implementing access controls and permissions to restrict unauthorized access to sensitive data.\nData Privacy: Ensuring data privacy and compliance with regulations such as GDPR and HIPAA.\nContinuous Integration and Continuous Deployment (CI/CD) Streamlining the deployment process through CI/CD pipelines is crucial for efficient and reliable production workflows. We\u0026rsquo;ll cover:\nCI/CD Pipelines: Implementing CI/CD pipelines for automated testing and deployment of Dockerized data science applications.\nContainer Registry: Setting up a private container registry to store and manage Docker images securely.\nRolling Updates: Strategies for performing rolling updates and rollbacks in a production environment.\nCase Studies: Docker in Production To provide real-world insights, we\u0026rsquo;ll examine case studies of organizations that have successfully deployed Dockerized data science applications in production. These case studies will illustrate how Docker can be used effectively to meet the demands of production environments while delivering value to businesses.\nBy the end of this chapter, you\u0026rsquo;ll have a comprehensive understanding of how to use Docker in a production environment for data science applications. You\u0026rsquo;ll be well-prepared to navigate the challenges of productionizing data science workflows, from infrastructure automation to security considerations and CI/CD integration. It\u0026rsquo;s time to take your Docker skills to the next level and confidently deploy data science solutions in real-world scenarios.\nChapter 8: Docker for Data Science: Future Trends and Beyond As the field of data science continues to evolve, so too does the role of Docker in enabling efficient, reproducible, and scalable data science workflows. In this final chapter, we\u0026rsquo;ll explore the future trends and emerging possibilities for Docker in the data science landscape, as well as provide guidance on staying up-to-date with the latest developments.\nThe Evolution of Docker for Data Science Docker has come a long way in the data science domain, and its role is expected to grow even more prominent in the future. We\u0026rsquo;ll discuss:\nKubernetes Integration: How Docker integrates with Kubernetes and other orchestration platforms to simplify container management and scaling. Docker and Kubernetes are a powerful combination for container orchestration. Here\u0026rsquo;s how to leverage this integration for data science:\n# Deploy a data science application in a Kubernetes cluster kubectl apply -f your_data_science_deployment.yaml Serverless Computing: The rise of serverless computing and how Docker containers fit into serverless architectures for data science applications. Serverless architectures are becoming more prevalent for data science tasks. Here\u0026rsquo;s how Docker fits into this landscape: Define a serverless function using Docker in a cloud provider - sample of yaml file functions: data-processing: image: your_data_processing_container:latest Cloud-Native Ecosystem: The convergence of Docker with cloud-native technologies, enabling data scientists to leverage cloud services seamlessly. The cloud-native ecosystem is expanding, offering data scientists more resources than ever. Learn how to incorporate Docker into your cloud-native data science workflows: # Use Docker for creating cloud-native microservices docker build -t your_microservice_image . docker run -d -p 8080:8080 your_microservice_image Edge Computing and Docker Edge computing is gaining traction as a critical component of data science, particularly for real-time and IoT applications. Edge computing is crucial for real-time and IoT data science applications. Here\u0026rsquo;s how Docker can be utilized at the edge: We\u0026rsquo;ll explore:\nEdge Deployment: How Docker containers can be deployed at the edge to perform data preprocessing, analytics, and machine learning inference. Deploying Docker containers at the edge for data preprocessing and analytics: # Deploy a Docker container at the edge for real-time data processing docker run -d --name edge-processor your_edge_container:latest Challenges and Opportunities: The challenges of edge computing in data science and how Docker addresses these challenges. Learn about the challenges and opportunities of using Docker at the edge in data science: - **Challenge:** Ensuring low latency and high availability in edge deployments. - **Opportunity:** Leveraging Docker\u0026#39;s lightweight containers for resource-efficient edge computing. Edge Use Cases: Real-world use cases of Docker at the edge in data science applications. Beyond Containers: Exploring Data Science with Containerd and CRI-O While Docker has been a dominant containerization platform, alternative runtimes such as Containerd and CRI-O are emerging. Docker has been the dominant containerization platform, but alternative runtimes like Containerd and CRI-O are emerging. Here\u0026rsquo;s how to explore these runtimes and maintain compatibility: We\u0026rsquo;ll examine:\nContainer Runtimes: An overview of Containerd and CRI-O and their potential impact on the data science containerization landscape. Get familiar with Containerd and CRI-O and their potential impact on the data science containerization landscape: # Using Containerd to run a data science workload containerd run -d --name data-workload your_data_container:latest Compatibility and Migration: Considerations for transitioning from Docker to alternative runtimes while preserving your data science workflows. Ensure a smooth transition from Docker to alternative runtimes, preserving your data science workflows: - **Compatibility:** Check for compatibility with your existing Docker containers and images. - **Migration:** Develop a migration plan to shift from Docker to the chosen runtime. Staying Current with Docker and Data Science To remain at the forefront of data science and Docker advancements, it\u0026rsquo;s essential to stay informed and adaptable. We\u0026rsquo;ll provide guidance on:\nLearning Resources: Where to find the latest tutorials, courses, and documentation for Docker and data science.\nCommunity Engagement: Engaging with the Docker and data science communities, including forums, conferences, and open-source projects.\nExperimentation and Innovation: Encouraging a culture of experimentation and innovation within your data science team to explore new possibilities with Docker.\nBy the end of this chapter, you\u0026rsquo;ll have a glimpse into the exciting future of Docker in data science and its potential impact on the evolving data science landscape. As you continue your journey in data science, embracing the latest trends and technologies, including Docker, will be key to staying competitive and pushing the boundaries of what\u0026rsquo;s possible in this dynamic field.\nChapter 9: Case Studies: Docker in Data Science In this chapter, we\u0026rsquo;ll delve into real-world case studies that demonstrate the tangible benefits and practical applications of Docker in data science projects. These case studies showcase how Docker addresses specific data science challenges, streamlines workflows, and empowers data scientists to achieve remarkable results through hands-on examples.\nCase Study 1: Reproducible Machine Learning with Docker Problem: A machine learning team faces issues with reproducibility and environment consistency. Each team member uses a slightly different setup, leading to inconsistencies in model results. A machine learning team faces reproducibility challenges due to inconsistent development environments.\nSolution: The team adopts Docker to containerize their machine learning environments. They create Docker images containing the necessary libraries, dependencies, and Jupyter notebooks for experimentation. Each team member uses the same Docker image, ensuring consistent environments across the team. The team adopts Docker to create reproducible environments. They define a Dockerfile as follows:\n# Use a base image with required dependencies FROM python:3.8 # Set the working directory WORKDIR /app # Install necessary packages RUN pip install pandas scikit-learn # Copy code into the container COPY . /app # Specify the command to run when the container starts CMD [\u0026#34;python\u0026#34;, \u0026#34;train_model.py\u0026#34;] Outcome: Docker ensures consistent environments across the team. Team members build the Docker image with docker build -t ml-environment . and run it with docker run ml-environment. This leads to better collaboration and more consistent model results. With Docker, the team achieves greater reproducibility, easier collaboration, and the ability to share pre-configured environments with external collaborators. Model results become more consistent, and onboarding new team members is simplified.\nCase Study 2: Scalable Data Processing with Docker and Kubernetes Problem: An e-commerce company must scale its data processing pipeline to handle growing data volumes. An e-commerce company processes vast amounts of user data and faces challenges in scaling their data processing pipeline to meet growing demands.\nSolution: They adopt a microservices architecture with Docker containers and leverage Kubernetes for orchestration. Data processing tasks are split into containers, and Kubernetes scales them based on demand. The company employs Docker for containerization and Kubernetes for orchestration. They create Kubernetes Deployment YAML files like this:\napiVersion: apps/v1 kind: Deployment metadata: name: data-processor spec: replicas: 3 template: metadata: labels: app: data-processor spec: containers: - name: data-processor image: your-data-processor-image:latest Outcome: With Docker and Kubernetes, the company achieves the scalability needed to process large datasets efficiently. Their data pipeline can adapt to traffic spikes and handle increased workloads, ensuring a smooth user experience. Docker containers run efficiently on Kubernetes, allowing the company to scale data processing tasks. Kubernetes scales the number of replicas based on demand, ensuring smooth data processing as traffic fluctuates.\nCase Study 3: Real-time Analytics on Edge Devices Problem: An IoT device manufacturer needs to perform real-time analytics on edge devices in remote locations. A company that manufactures IoT devices needs to perform real-time analytics on edge devices in remote locations with limited connectivity.\nSolution: They develop a Docker container that runs on edge devices, capturing and processing data locally. Docker\u0026rsquo;s lightweight and portable nature allows the same container to be used across various edge devices. They develop a lightweight Docker container that captures and processes data on edge devices. The Dockerfile looks like this:\n# Use a minimal base image FROM alpine:latest # Install necessary tools RUN apk --no-cache add your-tools # Copy and set up data processing scripts COPY data-processing.sh /app/ CMD [\u0026#34;sh\u0026#34;, \u0026#34;/app/data-processing.sh\u0026#34;] Outcome: By using Docker at the edge, the company achieves real-time analytics capabilities, reduces data transfer costs, and ensures timely responses to critical events in remote locations. Docker\u0026rsquo;s portability and efficiency allow the same container to run on various edge devices, enabling real-time analytics, reducing data transfer costs, and ensuring timely responses to events in remote locations.\nCase Study 4: Continuous Integration for Data Science Problem: A data science team frequently encounters integration issues when merging their code and models. They need a streamlined testing and deployment process. A data science team struggles with code and model integration issues when deploying new solutions.\nSolution: The team integrates Docker into their continuous integration (CI) workflow. They create Docker images that encapsulate their data processing pipelines and model training scripts, allowing for consistent testing in different environments. The team incorporates Docker into their continuous integration (CI) process. They create Docker images that encapsulate data processing pipelines and model training scripts, ensuring consistent testing in various environments. The .gitlab-ci.yml file includes:\nimage: docker:19.03 services: - docker:19.03-dind stages: - test - deploy variables: DOCKER_HOST: tcp://docker:2375 test: script: - docker build -t your-test-image . - docker run your-test-image python test.py deploy: script: - docker build -t your-production-image . - docker push your-production-image Outcome: By adopting Docker in their CI process, the team improves code and model quality. They can confidently deploy new models with minimal integration issues, resulting in more reliable data science solutions. With Docker in their CI process, the team achieves more reliable code and model integration. They confidently deploy new solutions with minimal issues.\nCase Study 5: Secure Healthcare Data Analysis Problem: A healthcare organization faces the challenge of securely analyzing sensitive patient data for research while complying with data privacy regulations. A healthcare organization needs to analyze sensitive patient data while complying with data privacy regulations.\nSolution: They use Docker to containerize their data analysis workflows. By implementing strong access controls and encryption within Docker containers, they ensure data security and compliance. The organization employs Docker to containerize data analysis workflows. Strong access controls and data encryption are implemented within Docker containers. The Docker Compose file for data analysis services includes:\nversion: \u0026#39;3\u0026#39; services: data-analysis: image: your-data-analysis-container:latest environment: - DATABASE_URL=your-secure-database-url secrets: - your-encryption-key Outcome: With Docker, the organization can conduct data analysis in a secure and compliant manner, accelerating research while protecting patient data. Docker ensures secure and compliant data analysis. By leveraging Docker, the organization accelerates research while safeguarding patient data.\nThese case studies highlight the versatility of Docker in data science, from improving reproducibility and scalability to enabling real-time analytics and ensuring data security. By examining these real-world examples, you\u0026rsquo;ll gain a deeper understanding of how Docker can be applied to solve diverse data science challenges.\nChapter 10: Docker and Dataloaders in Data Science In this chapter, we\u0026rsquo;ll explore the synergy between Docker containers and dataloaders in the context of data science and machine learning. Docker has revolutionized the way we manage and deploy environments, and dataloaders play a crucial role in handling datasets for model training. This chapter delves into how Docker and dataloaders can be integrated to streamline data science workflows and ensure reproducibility.\nLeveraging Dataloaders in Docker Containers Dataloaders are an essential component of machine learning workflows, enabling efficient data loading, batching, and preprocessing. When combined with Docker containers, dataloaders offer powerful advantages:\nData Transformation in Containers: We\u0026rsquo;ll explore how to define data transformations within Docker containers, allowing you to preprocess data directly within the container environment. When you\u0026rsquo;re working with machine learning projects, data preprocessing is often a critical step. Docker containers allow you to define and encapsulate data transformation steps within the container environment. This ensures that data preprocessing is consistent and can be easily reproduced across different computing environments. Example:\nSuppose you\u0026rsquo;re building a Docker container for an image classification task, and you need to resize images to a specific size before feeding them to your model. You can define this data transformation within your Dockerfile:\n# Use a base image with necessary libraries FROM tensorflow/tensorflow:latest # Set the working directory WORKDIR /app # Install any additional dependencies RUN pip install opencv-python # Copy your preprocessing script into the container COPY preprocess.py /app # Define the command to run when the container starts CMD [\u0026#34;python\u0026#34;, \u0026#34;preprocess.py\u0026#34;] In this example, the preprocess.py script could contain code to resize and preprocess images using the OpenCV library. By encapsulating this data transformation within the Docker container, you ensure that the preprocessing steps are consistent and reproducible across different systems.\nParallel Data Loading: Docker containers can be used to parallelize data loading and preprocessing tasks, which is especially beneficial when working with large datasets. Docker containers can be used to parallelize data loading and preprocessing tasks. This is particularly beneficial when dealing with large datasets, as you can spawn multiple containers, each responsible for loading and preprocessing a subset of the data. Parallel data loading can significantly speed up the data preparation phase.\nCode Sample:\nSuppose you have a large dataset of images to process. You can create multiple Docker containers, each responsible for loading and preprocessing a portion of the dataset in parallel. Here\u0026rsquo;s an example using Python\u0026rsquo;s multiprocessing:\nfrom multiprocessing import Process def preprocess_data(start_idx, end_idx): # Load and preprocess data from start_idx to end_idx # This function runs in a separate Docker container # Split the dataset into segments for parallel processing segments = [(0, 1000), (1001, 2000), (2001, 3000)] # Create Docker containers for each segment processes = [] for start, end in segments: p = Process(target=preprocess_data, args=(start, end)) p.start() processes.append(p) # Wait for all processes to finish for p in processes: p.join() In this code sample, each Docker container is responsible for processing a segment of the dataset in parallel. This approach can significantly reduce data preprocessing time.\nScaling Data Loading: You\u0026rsquo;ll learn how to scale data loading tasks across multiple Docker containers to handle even the most extensive datasets efficiently. When working with extremely large datasets that do not fit into memory, you can use Docker containers to scale data loading tasks efficiently. By distributing the data loading across multiple containers, you can take advantage of the resources available on the host machine, making it easier to manage and load massive datasets.\nCode Sample:\nSuppose you have a dataset that is too large to fit into memory on a single machine. You can use Docker containers to load and process different segments of the dataset concurrently. Here\u0026rsquo;s an example using Python\u0026rsquo;s multiprocessing:\nfrom multiprocessing import Process def load_and_process_data(segment): # Load and process data for a specific segment # This function runs in a separate Docker container # Define the dataset segments to be processed segments = [\u0026#34;segment1\u0026#34;, \u0026#34;segment2\u0026#34;, \u0026#34;segment3\u0026#34;] # Create Docker containers for each segment processes = [] for segment in segments: p = Process(target=load_and_process_data, args=(segment,)) p.start() processes.append(p) # Wait for all processes to finish for p in processes: p.join() In this code sample, each Docker container loads and processes a specific segment of the dataset. By running these containers in parallel, you can efficiently handle datasets that are too large to fit into memory on a single machine.\nLeveraging dataloaders within Docker containers enhances the scalability, parallelism, and consistency of data preprocessing in your machine learning workflows. This approach is particularly valuable when dealing with large datasets and complex preprocessing tasks. This explanation and code sample demonstrate how dataloaders can be effectively used within Docker containers for data preprocessing, parallel data loading, and scaling data loading tasks in machine learning workflows.\nCase Study: Reproducible Model Training with Docker and Dataloaders To illustrate the concepts discussed in this chapter, we\u0026rsquo;ll walk through a case study. We\u0026rsquo;ll build a Docker image for a machine learning environment, implement dataloaders for a specific dataset, and demonstrate how to train a model with full reproducibility.\nThis case study will include practical examples, Dockerfiles, and Python code snippets to help you understand how Docker and dataloaders can work together seamlessly in a data science project.\nBy the end of this chapter, you\u0026rsquo;ll have a deeper understanding of the advantages of integrating Docker containers and dataloaders in data science workflows. You\u0026rsquo;ll be equipped to create reproducible environments, scale data loading tasks efficiently, and streamline your machine learning projects.\nProblem Statement: Imagine you\u0026rsquo;re working on a computer vision project that involves training a deep learning model for image classification. Ensuring reproducibility across different environments is crucial for collaboration and model deployment. Docker and dataloaders will be employed to address this challenge.\nSolution: We\u0026rsquo;ll create a Docker image that encapsulates the entire machine learning environment, including dependencies, and use dataloaders to handle image loading and preprocessing within the Docker container.\nStep 1: Dockerfile for Reproducible Environment First, let\u0026rsquo;s create a Dockerfile that specifies the environment for our machine learning project:\n# Use a base image with necessary libraries FROM tensorflow/tensorflow:latest # Set the working directory WORKDIR /app # Copy requirements.txt and install dependencies COPY requirements.txt /app/ RUN pip install --no-cache-dir -r requirements.txt # Copy the entire project into the container COPY . /app In this example, the requirements.txt file contains a list of Python libraries required for the project. The COPY commands ensure that all project files are copied into the container.\nChapter 11: Optimizing Dockerized Machine Learning Workflows In this chapter, we delve into strategies and best practices for optimizing machine learning workflows that leverage Docker containers. While Docker provides an excellent solution for encapsulating environments and ensuring reproducibility, optimizing the workflows within these containers is essential for achieving efficiency, scalability, and better performance. We will explore techniques to streamline model training, manage resources effectively, and enhance the overall productivity of Dockerized machine learning projects.\nStreamlining Model Training in Docker Containers One of the key challenges in machine learning workflows is optimizing model training processes within Docker containers. In this section, we\u0026rsquo;ll cover:\nPersistent Model Storage: Learn how to persist trained models outside the Docker container, enabling easy model retrieval and deployment.\nIncremental Training: Explore strategies for implementing incremental training within Docker containers to efficiently update models with new data.\nExperiment Tracking: Integrate tools for experiment tracking within Docker containers to monitor and manage different model training runs effectively.\nIn the section \u0026ldquo;Streamlining Model Training in Docker Containers,\u0026rdquo; we\u0026rsquo;ll explore strategies and best practices to optimize the model training process within Docker containers. Streamlining this aspect is crucial for efficiency, reproducibility, and effective management of machine learning projects. Let\u0026rsquo;s delve into the key components:\nPersistent Model Storage Problem:\nWhen training machine learning models within Docker containers, it\u0026rsquo;s essential to address the challenge of persisting trained models beyond the container\u0026rsquo;s lifespan. Without a mechanism for persistent storage, valuable trained models could be lost when the container is stopped or removed.\nSolution:\nTo overcome this challenge, it\u0026rsquo;s advisable to decouple the model training process from model storage within the container. Use external storage solutions or cloud services to store trained models persistently.\nCode Example:\n# Save the trained model to an external location model.save(\u0026#39;/path/to/persistent/storage/my_model.h5\u0026#39;) Incremental Training Problem:\nIn real-world scenarios, models often need to be updated with new data. Running the entire training process from scratch every time new data arrives can be time-consuming and resource-intensive.\nSolution:\nImplement incremental training strategies that allow models to be updated with new data efficiently. This involves saving the existing model\u0026rsquo;s weights, loading them for further training, and incorporating new data.\nCode Example:\n# Load the pre-trained model model = load_model(\u0026#39;/path/to/persistent/storage/my_model.h5\u0026#39;) # Continue training with new data model.fit(new_data, epochs=5) Experiment Tracking Problem:\nKeeping track of experiments, hyperparameters, and performance metrics is crucial for effective model management. Within Docker containers, monitoring and logging experiments become more challenging.\nSolution:\nIntegrate experiment tracking tools such as TensorBoard or MLflow within Docker containers to log and visualize metrics. These tools help manage different model training runs and provide insights into model performance.\nCode Example:\n# Integrate TensorBoard for experiment tracking tensorboard_callback = tf.keras.callbacks.TensorBoard(log_dir=\u0026#39;/path/to/tensorboard/logs\u0026#39;, histogram_freq=1) model.fit(train_dataset, epochs=10, callbacks=[tensorboard_callback]) By implementing these strategies, you can streamline the model training process within Docker containers, ensuring that trained models are persisted, updates can be performed incrementally, and experiments are well-tracked and monitored. This enhances the reproducibility and efficiency of your machine learning workflows\nManaging Resources Effectively Efficient resource management is crucial for running machine learning workloads at scale. We\u0026rsquo;ll discuss techniques to manage resources effectively within Docker containers:\nContainer Orchestration: Explore container orchestration tools like Kubernetes to manage and scale Docker containers in a distributed environment.\nGPU Acceleration: Understand how to leverage GPU acceleration within Docker containers for faster model training.\nDynamic Resource Allocation: Implement dynamic resource allocation strategies to adapt to varying workloads and optimize resource utilization.\nEnhancing Productivity with Docker Compose Docker Compose is a powerful tool for defining and managing multi-container Docker applications. We\u0026rsquo;ll explore how to enhance productivity using Docker Compose in the context of machine learning:\nMulti-Service Architecture: Design multi-service architectures for machine learning projects using Docker Compose to manage interconnected components.\nEnvironment Configuration: Leverage Docker Compose for simplified environment configuration, allowing seamless collaboration and deployment.\nCase Study: Scalable Training with Kubernetes and Docker To illustrate the concepts discussed in this chapter, we\u0026rsquo;ll walk through a case study on optimizing model training with Kubernetes and Docker. This case study will provide hands-on examples, including Kubernetes configurations and Dockerfiles, showcasing how to achieve scalability and efficiency in a real-world machine learning project.\nBy the end of this chapter, you\u0026rsquo;ll be equipped with the knowledge and tools to optimize your Dockerized machine learning workflows, ensuring that your projects are not only reproducible but also efficient, scalable, and well-managed.\nPersistent Model Storage Problem: When training machine learning models within Docker containers, it\u0026rsquo;s essential to address the challenge of persisting trained models beyond the container\u0026rsquo;s lifespan. Without a mechanism for persistent storage, valuable trained models could be lost when the container is stopped or removed.\nSolution: To overcome this challenge, it\u0026rsquo;s advisable to decouple the model training process from model storage within the container. Use external storage solutions or cloud services to store trained models persistently.\nCode Example: # Save the trained model to an external location model.save(\u0026#39;/path/to/persistent/storage/my_model.h5\u0026#39;) Incremental Training Problem: In real-world scenarios, models often need to be updated with new data. Running the entire training process from scratch every time new data arrives can be time-consuming and resource-intensive.\nSolution: Implement incremental training strategies that allow models to be updated with new data efficiently. This involves saving the existing model\u0026rsquo;s weights, loading them for further training, and incorporating new data.\nCode Example: # Load the pre-trained model model = load_model(\u0026#39;/path/to/persistent/storage/my_model.h5\u0026#39;) # Continue training with new data model.fit(new_data, epochs=5) Experiment Tracking Problem: Keeping track of experiments, hyperparameters, and performance metrics is crucial for effective model management. Within Docker containers, monitoring and logging experiments become more challenging.\nSolution: Integrate experiment tracking tools such as TensorBoard or MLflow within Docker containers to log and visualize metrics. These tools help manage different model training runs and provide insights into model performance.\nCode Example: # Integrate TensorBoard for experiment tracking tensorboard_callback = tf.keras.callbacks.TensorBoard(log_dir=\u0026#39;/path/to/tensorboard/logs\u0026#39;, histogram_freq=1) model.fit(train_dataset, epochs=10, callbacks=[tensorboard_callback]) By implementing these strategies, you can streamline the model training process within Docker containers, ensuring that trained models are persisted, updates can be performed incrementally, and experiments are well-tracked and monitored. This enhances the reproducibility and efficiency of your machine learning workflows.\nManaging Resources Effectively In the section \u0026ldquo;Managing Resources Effectively,\u0026rdquo; we\u0026rsquo;ll discuss strategies to optimize resource utilization within Docker containers for machine learning workloads.\nContainer Orchestration Problem: Scaling machine learning workloads across multiple Docker containers in a distributed environment can be complex and challenging to manage manually.\nSolution: Utilize container orchestration tools like Kubernetes to automate the deployment, scaling, and management of Docker containers. Kubernetes provides features for load balancing, auto-scaling, and fault tolerance, making it ideal for managing machine learning workloads at scale.\nExample: apiVersion: apps/v1 kind: Deployment metadata: name: my-model spec: replicas: 3 selector: matchLabels: app: my-model template: metadata: labels: app: my-model spec: containers: - name: my-model-container image: my-model-image:latest resources: limits: cpu: \u0026#34;2\u0026#34; memory: \u0026#34;4Gi\u0026#34; GPU Acceleration Problem: Training deep learning models often requires significant computational resources, particularly GPU acceleration. Leveraging GPUs within Docker containers can be challenging.\nSolution: Utilize GPU-enabled Docker images and configure Docker containers to access GPU resources. Tools like NVIDIA Docker enable seamless integration of GPUs with Docker containers, allowing for faster model training.\nExample: docker run --gpus all my-gpu-enabled-image:latest Dynamic Resource Allocation Problem: Machine learning workloads often vary in resource requirements over time, leading to inefficient resource allocation within Docker containers.\nSolution: Implement dynamic resource allocation strategies within Docker containers to adapt to varying workload demands. This can involve auto-scaling resources based on workload metrics or dynamically adjusting resource limits.\nExample: import docker client = docker.from_env() container = client.containers.run(\u0026#39;my-model-image:latest\u0026#39;, detach=True, cpu_shares=512) container.update(cpu_shares=1024) # Increase CPU shares dynamically Conclusion By implementing these resource management strategies within Docker containers for machine learning workloads, you can optimize resource utilization, improve scalability, and enhance overall efficiency. This ensures that your machine learning projects can handle varying workloads effectively and make the most out of available resources.\n","permalink":"https://naeem-bebit.github.io/docker/","summary":"An introduction to Docker in Data Science","title":"Docker and Kubernetes"},{"content":"There are two popular modules for unit test Pythin which are unittest (prebuild with Python package) and pytest\nUnit test Inside the unit test there are several packages such mock\nPytest For more information about the Pytest, please refer to the link\npytest test_{name of the file}.py options such verbose to display the percentage of the test coverage\nPlease refer to the link for more information on the unitest_mock https://docs.python.org/3/library/unittest.mock.html\nMost unit test follows this pattern\nArrange - declare the arguments Act - declare the function Assert - equal to the intended output https://www.youtube.com/watch?v=sCthIEOaMI8\u0026amp;list=PLJsmaNFr5mNqSeuNepT3IaMrgzRMm9lQR\u0026amp;index=2 ","permalink":"https://naeem-bebit.github.io/unitest/","summary":"Why unit test for Python","title":"Python Unit Test"},{"content":"WIP Bash echo print \u0026amp;\u0026amp; and ; both command in the same line will || or exit codes $ last commands ./ current directory Ex.\necho \u0026#34;Hello world\u0026#34; \u0026gt; Hello world echo $? \u0026gt; 0 #exit codes successful \u0026gt; 1 #exit codes not successful anything is not 0 Command substitution\nvar_1=$(\u0026#34;Test\u0026#34;) # no space var_2=${Test2} echo \u0026#34;test $var_1\u0026#34; echo \u0026#34;test $var_2\u0026#34; Trivia cat lorem_ipsum Pipe Symbol |\nFunction\nfunction name_of_function(){ return } Give permission to the file chmod +x filename.sh ./filename.sh #type this in the command\ngrep is used for search can be used with the pipe command\n-c count command -v not in -r recursive -C context -A after -B before -rni recursive, line number time to count\n","permalink":"https://naeem-bebit.github.io/bash/","summary":"\u003ch2 id=\"wip-bash\"\u003eWIP Bash\u003c/h2\u003e\n\u003cp\u003e\u003ccode\u003eecho\u003c/code\u003e print\n\u003ccode\u003e\u0026amp;\u0026amp;\u003c/code\u003e and\n\u003ccode\u003e;\u003c/code\u003e both command in the same line will\n\u003ccode\u003e||\u003c/code\u003e or\nexit codes\n\u003ccode\u003e$\u003c/code\u003e last commands\n\u003ccode\u003e./\u003c/code\u003e current directory\nEx.\u003c/p\u003e\n\u003cdiv class=\"highlight\"\u003e\u003cpre tabindex=\"0\" style=\"color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;\"\u003e\u003ccode class=\"language-bash\" data-lang=\"bash\"\u003e\u003cspan style=\"display:flex;\"\u003e\u003cspan\u003eecho \u003cspan style=\"color:#e6db74\"\u003e\u0026#34;Hello world\u0026#34;\u003c/span\u003e\n\u003c/span\u003e\u003c/span\u003e\u003cspan style=\"display:flex;\"\u003e\u003cspan\u003e\u0026gt; Hello world\n\u003c/span\u003e\u003c/span\u003e\u003cspan style=\"display:flex;\"\u003e\u003cspan\u003e\n\u003c/span\u003e\u003c/span\u003e\u003cspan style=\"display:flex;\"\u003e\u003cspan\u003eecho $?\n\u003c/span\u003e\u003c/span\u003e\u003cspan style=\"display:flex;\"\u003e\u003cspan\u003e\u0026gt; \u003cspan style=\"color:#ae81ff\"\u003e0\u003c/span\u003e \u003cspan style=\"color:#75715e\"\u003e#exit codes successful\u003c/span\u003e\n\u003c/span\u003e\u003c/span\u003e\u003cspan style=\"display:flex;\"\u003e\u003cspan\u003e\u0026gt; \u003cspan style=\"color:#ae81ff\"\u003e1\u003c/span\u003e \u003cspan style=\"color:#75715e\"\u003e#exit codes not successful anything is not 0\u003c/span\u003e\n\u003c/span\u003e\u003c/span\u003e\u003c/code\u003e\u003c/pre\u003e\u003c/div\u003e\u003cp\u003eCommand substitution\u003c/p\u003e\n\u003cdiv class=\"highlight\"\u003e\u003cpre tabindex=\"0\" style=\"color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;\"\u003e\u003ccode class=\"language-bash\" data-lang=\"bash\"\u003e\u003cspan style=\"display:flex;\"\u003e\u003cspan\u003evar_1\u003cspan style=\"color:#f92672\"\u003e=\u003c/span\u003e\u003cspan style=\"color:#66d9ef\"\u003e$(\u003c/span\u003e\u003cspan style=\"color:#e6db74\"\u003e\u0026#34;Test\u0026#34;\u003c/span\u003e\u003cspan style=\"color:#66d9ef\"\u003e)\u003c/span\u003e \u003cspan style=\"color:#75715e\"\u003e# no space\u003c/span\u003e\n\u003c/span\u003e\u003c/span\u003e\u003cspan style=\"display:flex;\"\u003e\u003cspan\u003evar_2\u003cspan style=\"color:#f92672\"\u003e=\u003c/span\u003e\u003cspan style=\"color:#e6db74\"\u003e${\u003c/span\u003eTest2\u003cspan style=\"color:#e6db74\"\u003e}\u003c/span\u003e\n\u003c/span\u003e\u003c/span\u003e\u003cspan style=\"display:flex;\"\u003e\u003cspan\u003eecho \u003cspan style=\"color:#e6db74\"\u003e\u0026#34;test \u003c/span\u003e$var_1\u003cspan style=\"color:#e6db74\"\u003e\u0026#34;\u003c/span\u003e\n\u003c/span\u003e\u003c/span\u003e\u003cspan style=\"display:flex;\"\u003e\u003cspan\u003eecho \u003cspan style=\"color:#e6db74\"\u003e\u0026#34;test \u003c/span\u003e$var_2\u003cspan style=\"color:#e6db74\"\u003e\u0026#34;\u003c/span\u003e\n\u003c/span\u003e\u003c/span\u003e\u003c/code\u003e\u003c/pre\u003e\u003c/div\u003e\u003ch2 id=\"trivia\"\u003eTrivia\u003c/h2\u003e\n\u003cdiv class=\"highlight\"\u003e\u003cpre tabindex=\"0\" style=\"color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;\"\u003e\u003ccode class=\"language-bash\" data-lang=\"bash\"\u003e\u003cspan style=\"display:flex;\"\u003e\u003cspan\u003ecat lorem_ipsum\n\u003c/span\u003e\u003c/span\u003e\u003c/code\u003e\u003c/pre\u003e\u003c/div\u003e\u003cp\u003ePipe\nSymbol \u003ccode\u003e|\u003c/code\u003e\u003c/p\u003e\n\u003cp\u003eFunction\u003c/p\u003e\n\u003cdiv class=\"highlight\"\u003e\u003cpre tabindex=\"0\" style=\"color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;\"\u003e\u003ccode class=\"language-bash\" data-lang=\"bash\"\u003e\u003cspan style=\"display:flex;\"\u003e\u003cspan\u003e\u003cspan style=\"color:#66d9ef\"\u003efunction\u003c/span\u003e name_of_function\u003cspan style=\"color:#f92672\"\u003e(){\u003c/span\u003e\n\u003c/span\u003e\u003c/span\u003e\u003cspan style=\"display:flex;\"\u003e\u003cspan\u003e    \u003cspan style=\"color:#66d9ef\"\u003ereturn\u003c/span\u003e \n\u003c/span\u003e\u003c/span\u003e\u003cspan style=\"display:flex;\"\u003e\u003cspan\u003e\u003cspan style=\"color:#f92672\"\u003e}\u003c/span\u003e\n\u003c/span\u003e\u003c/span\u003e\u003c/code\u003e\u003c/pre\u003e\u003c/div\u003e\u003cp\u003eGive permission to the file\n\u003ccode\u003echmod +x filename.sh\u003c/code\u003e\n./filename.sh #type this in the command\u003c/p\u003e","title":"Bash"},{"content":"UML Unified Modeling Language wiki\n","permalink":"https://naeem-bebit.github.io/data/science/uml/","summary":"Unified Modeling Language","title":"UML"},{"content":" Table of Contents Major types of data Numerical Data Categorical Data Ordinal Data Data type Major types of data Numerical Categorical Ordinal Numerical Data Represent quantitative measurement example: heights - 1.6m, 2.0m Discrete data example: age - 20,50 Continuous data Infinite number of possible values example: Temperature sensor Categorical Data The data doesn\u0026rsquo;t have mathematical meaning example: gender, race Ordinal Data A mixture of categorical and numerical data example: movie rating - one star is low rating compared to 5 star rating ","permalink":"https://naeem-bebit.github.io/data_types/","summary":"Most common type of data","title":"Type of Data"},{"content":"There are a lot of web frameworks available to choose from. Yet there are only two web frameworks which are very popular among Python developers. There are Django and Flask. More details on other Python web frameworks can be check here.\nFlask VS Django Thus in the post, will present the comparison between the most popular web frameworks which are Flask and Django.\nThe biggest difference between Flask and Django is: Flask implements a bare-minimum and leaves the bells and whistles to add-ons or to the developer Django follows a \u0026ldquo;batteries included\u0026rdquo; philosophy and gives you a lot more out of the box. In the nutshell the comparison can be presented in analog like Flask is more like a pirates and Django is more like a Navy.\nDjango released in 2005 and Flask released in 2010 Flask How to install Flask\npip install flask Create a Python file called flaskhelloworld.py and insert the following code\nfrom flask import Flask app = Flask(__name__) @app.route(\u0026#34;/\u0026#34;) def hello(): return \u0026#34;Hello, World!\u0026#34; if __name__ == \u0026#34;__main__\u0026#34;: app.run() Run the command\npython3 flaskhelloworld.py Application is running on and \u0026lsquo;127.0.0.1\u0026rsquo; means that the application is running on local host — it\u0026rsquo;s only accessible on our development machine. If you open a web browser and visit http://127.0.0.1:5000/, you\u0026rsquo;ll see a web page that returns the \u0026ldquo;Hello, World!\u0026rdquo; greeting.\nDjango pip install django Set up the django-admin command. Run the following:\ndjango-admin startproject hellodjango python3 manage.py startapp helloworld Open helloworld/views.py and the following code:\nfrom django.http import HttpResponse def index(request): return HttpResponse(\u0026#34;Hello, World!\u0026#34;) From the hellodjango directory, run the following command:\npython3 manage.py startapp helloworld In the hellodjango directory too, open the automatically created helloworld/views.py file and add the following code:\nfrom django.http import HttpResponse def index(request): return HttpResponse(\u0026#34;Hello, World!\u0026#34;) Also need to create a urls.py file for the application. Create helloworld/urls.py and add the following code:\nfrom django.conf.urls import url from . import views urlpatterns = [ url(r\u0026#39;^$\u0026#39;, views.index, name=\u0026#39;index\u0026#39;), ] hellodjango directory (the one which contains the manage.py file) and run the following command:\npython3 manage.py runserver ","permalink":"https://naeem-bebit.github.io/webframeworks/backend/","summary":"Django VS Flask","title":"Web Frameworks for Python"},{"content":"Data Structures Different ways of storing data on the computer or more simple way is a database\nArray Collection of items in single type\narray cant be changed, once it has defined, it cannot be changed. For example in Python it\u0026rsquo;s called tuple\ntuple1 = (1,2,3) tuple1 cannot be changed since it\u0026rsquo;s immutable, thus in order to change the tuple it must be converted to the list first and add another element to the list via append and change it back to tuple. Once a tuple is created, you cannot add items to it. Tuples are unchangeable. Example:\nx = (\u0026#34;apple\u0026#34;, \u0026#34;banana\u0026#34;, \u0026#34;cherry\u0026#34;) y = list(x) y[1] = \u0026#34;kiwi\u0026#34; x = tuple(y) thislist = [\u0026#34;apple\u0026#34;, \u0026#34;banana\u0026#34;, \u0026#34;cherry\u0026#34;] thislist.append(\u0026#34;orange\u0026#34;) Example:\nthistuple = (\u0026#34;apple\u0026#34;, \u0026#34;banana\u0026#34;, \u0026#34;cherry\u0026#34;) thistuple[3] = \u0026#34;orange\u0026#34; # This will raise an error Memory a long tape of bytes (byte is a small unit of data consists of 8 bits) Ram - temporary memory Disk - permenant memory\nLinked List Built-in Data Structures List Dictionary Set Tuple User-Defined Data Structures Stack Queue Tree Linked List Graph HashMap many more (build mindmap using the data structures) https://www.youtube.com/watch?v=xOuRE3IuEB8 Algorithm Operations on different data structures and sets of instructions for executing it or simplay way to define it is programming\n","permalink":"https://naeem-bebit.github.io/algo/","summary":"An introduction to the essence of programming","title":"Data Structure and Algorithm"},{"content":"In this post, I would like to share about the useful Python libraries.\nPyforest Lazy-import of all popular Python Data Science libraries. Stop writing the same imports over and over again.\npip install pyforest or\npip install --upgrade pyforest It supports almost most of the popular libraries such as pd(pandas), np(numpy), sklearn etc. You can check the supported libraries by running below command\ndir(pyforest) In the jupyter notebook just run command below and you dont have to import for each libraries when you want to use it.\nimport pyforest To check what libraries you have imported you can run command below\nactive_imports() Then you dont have to import all the libraries at the beginning each time, so give it a try!\nProphet How to install it\npip install fbprophet python import\nfrom fbprophet import Prophet df = pd.read_csv(\u0026#39;data.csv\u0026#39;) m = Prophet() m.fit(df) #The predict method will assign each row in future a predicted value forecast = m.predict(future) forecast[[\u0026#39;ds\u0026#39;, \u0026#39;yhat\u0026#39;, \u0026#39;yhat_lower\u0026#39;, \u0026#39;yhat_upper\u0026#39;]].tail() #Plot fig1 = m.plot(forecast) fig2 = m.plot_components(forecast) #It also has it owns plot from fbprophet.plot import plot_plotly import plotly.offline as py py.init_notebook_mode() fig = plot_plotly(m, forecast) # This returns a plotly Figure py.iplot(fig) Code Formatter Black Black is the uncompromising Python code formatter A simple extension for Jupyter Notebook and Jupyter Lab to beautify Python code automatically using Black\npip install nb_black For Jupyter Notebook\n%load_ext nb_black d = {\u0026#34;a\u0026#34;: 1, \u0026#34;b\u0026#34;:2, \u0026#39;c\u0026#39; : \u0026#34;nikhil\u0026#34;} Formatted\nd = {\u0026#34;a\u0026#34;: 1, \u0026#34;b\u0026#34;: 2, \u0026#34;c\u0026#34;: \u0026#34;nikhil\u0026#34;} Progress Bar in Jupyter notebook Library tqdm\npip install tqdm In the Jupyter notebook\nfrom tqdm import trange from time import sleep for i in trange(100): sleep(0.01) Code Refactoring Radon pip install radon Radon is a Python tool that computes various metrics from the source code. Radon can compute:\nMcCabe’s complexity, i.e. cyclomatic complexity raw metrics (these include SLOC, comment lines, blank lines, \u0026amp;c.) Halstead metrics (all of them) Maintainability Index (the one used in Visual Studio) Data model Special Methods under under method under under = Dunder __name__ == __main__: main() The function’s name.\nYou can refer for more information about Dunder Methods\nstreamlit import streamlit as st ## Title st.title(\u0026#34;Hello World\u0026#34;) ## Markdown st.markdown(\u0026#34;This example of Markdown\u0026#34;) ## Header/subheader st.header(\u0026#34;Header\u0026#34;) st.subheader(\u0026#34;Subheader\u0026#34;) ## Colourful(Error/Information) st.successful(\u0026#34;Successful\u0026#34;) st.warning(\u0026#34;Error\u0026#34;) st.info(\u0026#34;Information\u0026#34;) st.error(\u0026#34;Error\u0026#34;) st.exception(\u0026#34;Exception\u0026#34;) from PIL import Image img = Image.open(\u0026#34;example.jpg\u0026#34;) import datetime today = st.date_input(\u0026#34;Today is \u0026#34;, datetime.datetime.now) the_time = st.time_input(\u0026#34;The time is\u0026#34;,datetime.time()) if st.button(\u0026#34;Thanks\u0026#34;): st.balloons() st.sidebar.header(\u0026#34;About App\u0026#34;) st.sidebar.info(\u0026#34;A Visualization Demo for TRI Data Scientist Challenge\u0026#34;) st.sidebar.header(\u0026#34;About\u0026#34;) st.sidebar.info(\u0026#34;Naeem Hussien\u0026#34;) st.sidebar.text(\u0026#34;Machine Learning Engineer\\n\u0026#34; \u0026#34;beBit Tokyo Japan\u0026#34;) ","permalink":"https://naeem-bebit.github.io/py-lib/","summary":"Interesting \u0026amp; Useful Python Library","title":"Python Library"},{"content":"WIP (Work in Progress) CAPTURE, STORE, TRANSFORM, PUBLISH, and CONSUME\nCapture Persistent and resilient data CAPTURE is the first step in any big data system. Cloud Vendors and the community also describe data CAPTURE as ingest, extract, collect, or more generally as data movement. Data CAPTURE includes ingestion of both batch and streaming data.\nSTORE For big data systems the STORE stage focuses on the concept of a data lake, a single location where structured, semi-structured, unstructured data and objects are stored together. The data lake is also a place to store the output from extract, transform, load (ETL) and ML pipelines running in the TRANSFORM stage. Vendors focus on scalability and resilience over read/write performance. To increase data access and analytics performance, data should be highly aggregated in the data lake or organized and placed into higher performance data warehouses, massively parallel processing (MPP) databases, or key-value stores as described in the PUBLISH stage. In addition, some data streams have such high event volume, or the data are only relevant at the time of capture, that the data stream may be processed without ever entering the data lake.\nTRANSFORM The heart of any big data implementation is the ability to create data pipelines in order to clean, prepare, and TRANSFORM complex multi-modal data into valuable information. Data TRANSFORM is also described as preparing, massaging, processing, organizing, and analyzing among other things. The TRANSFORM stage is where value is created and, as a result, Cloud Vendors, start-ups, and traditional database and ETL vendors provide many tools. The TRANSFORM stage has three main data pipeline offerings including Batch Processing, Machine Learning, and Stream Processing. In addition, we include the Orchestration offering because complex data pipelines require tools to stage, schedule, and monitor deployments.\nML/AI uses many of the same Batch Processing tools and techniques for data preparation and for the development and training of predictive models. Machine Learning also takes advantage numerous libraries and packages to help optimize data science workflows and provide pre-built algorithms. Big data systems also provide tools to query continuous data streams in near real-time. Some data has immediate value that would be lost waiting for a batch process to run. For example, predictive models for fraud detection or alerts based on data from an IoT sensor. In addition, streaming data is commonly processed, and portions are loaded into a data lake.\nPUBLISH Once through the data CAPTURE and TRANSFORM stages it is necessary to PUBLISH the output from batch, ML, or streaming pipelines for users and applications to CONSUME. PUBLISH is also described as deliver or serve, and comes in the form of Data Warehouses, Data Catalogs, or Real-Time Stores.\nCONSUME The value of any big data system comes together in the hands of technical and non-technical users, and in the hands of customers using data-centric applications and products. Vendors also refer to CONSUME as use, harness, explore, model, infuse, and sandbox. We discuss three CONSUME models: Advanced Analytics, Business Intelligence (BI), and Real-Time APIs.\n","permalink":"https://naeem-bebit.github.io/machine%20learning/arch/","summary":"Data Architecture for Data Science","title":"Data Architecture and Data Science"},{"content":"WIP (Work in Progress) In this post, I will share some of the Python programming\nThe Zen of Python by Tim Peters\nimport this Beautiful is better than ugly. Explicit is better than implicit. Simple is better than complex. Complex is better than complicated. Flat is better than nested. Sparse is better than dense. Readability counts. Special cases aren\u0026rsquo;t special enough to break the rules. Although practicality beats purity. Errors should never pass silently. Unless explicitly silenced. In the face of ambiguity, refuse the temptation to guess. There should be one\u0026ndash; and preferably only one \u0026ndash;obvious way to do it. Although that way may not be obvious at first unless you\u0026rsquo;re Dutch. Now is better than never. Although never is often better than right now. If the implementation is hard to explain, it\u0026rsquo;s a bad idea. If the implementation is easy to explain, it may be a good idea. Namespaces are one honking great idea \u0026ndash; let\u0026rsquo;s do more of those! __main__ __init__ decorators Decorators is a function where we want to decorate a function without changing the original function\ndef dev(a,b): return a/b dev(2,4) Result\n0.5 Let say we want to execute above function and want to alter the function or add another into the original function, we use decorator\ndef smart_dev(func): def inner_function(a,b): if a\u0026lt;b: a,b = b,a return func(a,b) return inner_function @smart_dev def dev(a,b): return a/b dev(2,4) Result\n2 https://www.youtube.com/watch?v=FsAPt_9Bf3U\nClass and object Why Class in Python? To group the data and function for easy to use and build upon\nObject is the subset of the class. The object is created by the defined object Class Example:\nclass Class_Name: def __init__(self, att_1, att_2, att_3): self.att_1 = att_1 self.att_2 = att_2 self.att_3 = att_3 def function_class(self): print(self.att_1) Method Method is the function associated with the class Add self to each of the method of the class Class_Name\ndef function_class(self): Constructor The custome constructor like example below is how to define the attributes in more efficient way\ndef __init__(self, att_1, att_2, att_3): Therefore to run the example of the Class above\nvariable1 = Class_Name(att_1, att_2, att_3) String formatting There are multiple ways on how you can do string-formatting Let say we have an output like this\nMy name is Ahmad and I am 20 years old. Example\nperson = {\u0026#39;name\u0026#39;: \u0026#39;Ahmad\u0026#39;, \u0026#39;age\u0026#39;: 20} sentence = \u0026#39;My name is \u0026#39; + person[\u0026#39;name\u0026#39;] + \u0026#39; and I am \u0026#39; + str(person[\u0026#39;age\u0026#39;]) + \u0026#39; years old.\u0026#39; On the example above, you need to convert the integer to string due to error TypeError: can only concatenate str (not \u0026quot;int\u0026quot;) to str\nYou could refer to the date time format based on official documentation of Python datetime library\nimport datetime my_date = datetime.datetime(2016, 9, 24, 12, 30, 45) sentence = \u0026#39;{0:%B %d, %Y} fell on a {2:%d} and was the {1:%y} day of the year\u0026#39;.format(my_date, my_date, my_date) On the example above you can either add 0 at the beginning of the :%B %d, %Y or you put the 3 placeholders but provide only a value to the format. Therefore you can either put 3 values in the format or assign only one value to the format and add 0:.\nstr() and repr() str should be readable repr should be unambigious\nNamedtuple To increase the readability of the tuple\nfrom collections import namedtuple argparse Argparse is used to parse python scripts parameters\nimport argparse if __name__ = \u0026#39;__main__\u0026#39;: # Initialize parser parser = argparse.ArgumentParser() parser.add_argument(\u0026#39;\u0026#39;) # Add the additional parameter and add the parser args = parser.parse_args() ","permalink":"https://naeem-bebit.github.io/basic/","summary":"An introduction to Python programming","title":"Basic Python"},{"content":"WIP (Work in Progress) What is cloud computing? Cloud Computing is the on-demand availability of computer system resources, especially data storage and computing power, without direct active management by the user. The term is generally used to describe data centers available to many users over the Internet.\nIn the past five years, a shift in Cloud Vendor offerings has fundamentally changed how companies buy, deploy and run big data systems. Cloud Vendors have absorbed more back-end data storage and transformation technologies into their core offerings and are now highlighting their data pipeline, analysis, and modeling tools. This is great news for companies deploying, migrating, or upgrading big data systems. Companies can now focus on generating value from data and Machine Learning (ML), rather than building teams to support hardware, infrastructure, and application deployment/monitoring.\nAs the technology gets easier to deploy, and the Cloud Vendor data services mature, it becomes much easier to build data-centric applications and provide data and tools to the enterprise. This is good news: companies looking to migrate from on-premise systems to the cloud are no longer required to purchase directly or manage hardware, storage, networking, virtualization, applications, and databases. In addition, this changes the operational focus for a big data systems from infrastructure and application management (DevOps) to pipeline optimization and data governance (DataOps). The following table shows the different roles required to build and run Cloud Vendor-based big data systems.\nCloud Vendor Google Cloud Platform (GCP) Microsoft Azure Amazon AWS ","permalink":"https://naeem-bebit.github.io/cloud/aws/","summary":"Why using Cloud for Data Science","title":"Cloud Computing and Data Science"},{"content":"In this post, I would cover the function inside the library called pandas\nhttps://pandas.pydata.org/docs/reference/general_functions.html\n","permalink":"https://naeem-bebit.github.io/ml/","summary":"An introduction to Machine Learning","title":"Machine Learning"},{"content":"WIP (Work in Progress) What is Git? Git is a free and open source distributed version control system designed to handle everything from small to very large projects with speed and efficiency\nWhy data science need to use Git? Using Git in a group project allows for several developers to work on the same project independently without constantly interfering with each other’s input. Each developer can get an independent version of the code that they can modify without taking the risk of destroying a stable version of the code.\nThe ability to duplicate the code and work independently on a different version of it makes Git a good option for anyone building an application, even a developer working alone. It gives you the opportunity to keep several versions of your code and keep track of all characteristics of each change, such as who did the change and when.\nBasic command \u0026ndash; will add more explanation later git add git commit git push git pull git fetch git rebase git remote git merge Few examples of Git repository Github Gitlab BitBucket git submodules ","permalink":"https://naeem-bebit.github.io/git/","summary":"Why using Git for Data Science","title":"Git and Data Science"},{"content":"After obtaining a Bachelor\u0026rsquo;s Degree in Computer Engineering (CE), I soon pursued my Masters in Image Processing, which was a field initially considered as an Engineering discipline but later on found its relevance towards Computer Science.\nAt the time (circa 2013), the Data revolution was just beginning to take place and people were still introduced to the term ‘Big Data’ due to a need for a more efficient way of managing vast \u0026amp; complex data/information.\nIt is through Machine Learning that I have discovered my true calling; that is a career in Data Science. However, the transition to venture into this field takes a lot of effort and courage to get out of my comfort zone. Not to mention having to self-taught myself in Data Science due to its fairly recent arrival in a country like Malaysia.\nDespite having a considerable amount of Image Processing \u0026amp; Programming backgrounds, there were still challenges that I’ve encountered throughout my Data Science journey, such as having to revisit Statistics studies and journal-reading.\nEmbarking on this career path was not a breeze but it definitely provided me with a lot of invaluable experience so far. I hope that this blog on Data Science would benefit both the readers and myself as it’s not only just me sharing my views but it’s also about exchanging ideas and experiences with fellow programmers around the world.\nLet’s code to save the world!\n","permalink":"https://naeem-bebit.github.io/about/","summary":"\u003cp\u003eAfter obtaining a Bachelor\u0026rsquo;s Degree in Computer Engineering (CE), I soon pursued my Masters in Image Processing, which was a field initially considered as an Engineering discipline but later on found its relevance towards Computer Science.\u003c/p\u003e\n\u003cp\u003eAt the time (circa 2013), the Data revolution was just beginning to take place and people were still introduced to the term ‘Big Data’ due to a need for a more efficient way of managing vast \u0026amp; complex data/information.\u003c/p\u003e","title":"About Me"},{"content":"Feel free to ask me on data science/machine learning or image processing by using the contact form below.\nI will get back to you as soon as possible! :)\nYour Name: Your Email: Subject: Your Location: Message: Send ","permalink":"https://naeem-bebit.github.io/contact/","summary":"\u003cp\u003eFeel free to ask me on data science/machine learning or image processing by using the contact form below.\u003c/p\u003e\n\u003cp\u003eI will get back to you as soon as possible! :)\u003c/p\u003e\n\u003cform name=\"contact\" method=\"POST\" action=\"https://formspree.io/xayozngg\"\u003e\n  \u003cp\u003e\n    \u003clabel\u003eYour Name: \u003cinput type=\"text\" name=\"name\" /\u003e\u003c/label\u003e\n  \u003c/p\u003e\n  \u003cp\u003e\n    \u003clabel\u003eYour Email: \u003cinput type=\"email\" name=\"email\" /\u003e\u003c/label\u003e\n  \u003c/p\u003e\n  \u003cp\u003e\n    \u003clabel\u003eSubject: \u003cinput type=\"text\" name=\"subject\" /\u003e\u003c/label\u003e\n  \u003c/p\u003e\n  \u003cp\u003e\n    \u003clabel\u003eYour Location: \u003cinput type=\"text\" name=\"location\" /\u003e\u003c/label\u003e\n  \u003c/p\u003e\n  \u003cp\u003e\n    \u003clabel\u003eMessage: \u003ctextarea name=\"message\"\u003e\u003c/textarea\u003e\u003c/label\u003e\n  \u003c/p\u003e\n  \u003cp\u003e\n    \u003cbutton type=\"submit\"\u003eSend\u003c/button\u003e\n  \u003c/p\u003e\n    \u003cinput type=\"hidden\" name=\"_next\" value=\"https://naeem-bebit.github.io/thankyou.html\"/\u003e\n    \u003cinput type=\"text\" name=\"_gotcha\" style=\"display:none\" /\u003e\n\u003c/form\u003e","title":"Contact"},{"content":"Thank you for your email. I will get back to you as soon as possible. Feel free to read my other blog posts\n","permalink":"https://naeem-bebit.github.io/thankyou.html","summary":"\u003cp\u003eThank you for your email. I will get back to you as soon as possible. Feel free to read my other \u003ca href=\"/posts/\"\u003eblog posts\u003c/a\u003e\u003c/p\u003e","title":"Thank you email"}]