
۴۴٬۰۰۰تومان
نوع فایل دانلود: EPUB
پس از خرید، یک فایل EPUB دریافت میکنید.
این فایل با Calibre، Apple Books و سایر کتابخوانهای دیجیتال مناسب است.
خلاصه
داسک ابزاری بومی برای تحلیل موازی است که طراحی شده تا بیدردسر با کتابخانههایی که همین حالا استفاده میکنید، مثل Pandas، NumPy و Scikit-Learn، هماهنگ شود. با داسک میتوانید دادههای بسیار بزرگ را پردازش کنید و با همان ابزارهایی که میشناسید کار کنید. و «دادهکاوی با پایتون و داسک» راهنمای شما برای استفاده از داسک در پروژههای دادهمحور است؛ بدون اینکه شیوه کارتان را عوض کنید!
خرید نسخه چاپی این کتاب شامل یک ایبوک رایگان در قالبهای PDF، Kindle و ePub از انتشارات Manning Publications است. داخل نسخه چاپی، دستورالعملهای ثبتنام را پیدا خواهید کرد.
درباره فناوری
یک خط لوله دادهکارآمد، برای موفقیت هر پروژه دادهکاوی حیاتی است. داسک یک کتابخانه منعطف برای محاسبات موازی در پایتون است که ساخت جریانهای کاری شهودی برای دریافت و تحلیل دادههای بزرگ و توزیعشده را آسان میکند. داسک زمانبندی پویا برای وظایف و مجموعههای موازی ارائه میدهد که قابلیتهای NumPy، Pandas و Scikit-learn را گسترش میدهد. در نتیجه، کاربران میتوانند کد خود را بهراحتی از یک لپتاپ تکنفره تا یک خوشه متشکل از صدها ماشین مقیاس دهند.
درباره کتاب
«دادهکاوی با پایتون و داسک» به شما کمک میکند پروژههای مقیاسپذیری بسازید که بتوانند دادههای عظیم را مدیریت کنند. بعد از آشنا شدن با چارچوب داسک، دادهها را در پایگاه داده «NYC Parking Ticket» تحلیل میکنید و از DataFrameها برای روانتر کردن فرایندتان استفاده میکنید. سپس با Dask-ML مدلهای یادگیری ماشین میسازید، تجسمهای تعاملی ایجاد میکنید و با AWS و Docker خوشهها را راهاندازی میکنید.
آنچه درون کتاب میبینید
کار با دادههای بزرگ و ساختیافته و همچنین دادههای بدون ساختار
تجسم با Seaborn و Datashader
پیادهسازی الگوریتمهای خودتان
ساخت برنامههای توزیعشده با Dask Distributed
بستهبندی و استقرار برنامههای داسک
درباره خواننده
مناسب برای دانشمندان داده و توسعهدهندگانی که تجربه کار با پشته PyData و پایتون را دارند.
درباره نویسنده
جسی دنیل یک توسعهدهنده باتجربه پایتون است. او در دانشگاه دنور، پایتون را برای دادهکاوی آموزش داده و تیمی از دانشمندان داده را در یک شرکت فناوری رسانهای مستقر در دنور رهبری میکند.
فهرست مطالب
بخش ۱ - بلوکهای سازنده برای محاسبات مقیاسپذیر
چرا محاسبات مقیاسپذیر اهمیت دارد
معرفی داسک
بخش ۲ - کار با دادههای ساختیافته با Dask DataFrames
معرفی Dask DataFrames
بارگذاری دادهها در DataFrames
پاکسازی و تبدیل DataFrames
جمعبندی و تحلیل DataFrames
تصور دادههای DataFrame با Seaborn
تصور دادههای مکانی با Datashader
بخش ۳ - گسترش و استقرار داسک
کار با Bagها و آرایهها
یادگیری ماشین با Dask-ML
مقیاسبندی و استقرار داسک
SummaryDask is a native parallel analytics tool designed to integrate seamlessly with the libraries you're already using, including Pandas, NumPy, and Scikit-Learn. With Dask you can crunch and work with huge datasets, using the tools you already have. And Data Science with Python and Dask is your guide to using Dask for your data projects without changing the way you work!Purchase of the print book includes a free eBook in PDF, Kindle, and ePub formats from Manning Publications. You'll find registration instructions inside the print book.About the TechnologyAn efficient data pipeline means everything for the success of a data science project. Dask is a flexible library for parallel computing in Python that makes it easy to build intuitive workflows for ingesting and analyzing large, distributed datasets. Dask provides dynamic task scheduling and parallel collections that extend the functionality of NumPy, Pandas, and Scikit-learn, enabling users to scale their code from a single laptop to a cluster of hundreds of machines with ease.About the BookData Science with Python and Dask teaches you to build scalable projects that can handle massive datasets. After meeting the Dask framework, you'll analyze data in the NYC Parking Ticket database and use DataFrames to streamline your process. Then, you'll create machine learning models using Dask-ML, build interactive visualizations, and build clusters using AWS and Docker.What's insideWorking with large, structured and unstructured datasetsVisualization with Seaborn and DatashaderImplementing your own algorithmsBuilding distributed apps with Dask DistributedPackaging and deploying Dask appsAbout the ReaderFor data scientists and developers with experience using Python and the PyData stack.About the AuthorJesse Daniel is an experienced Python developer. He taught Python for Data Science at the University of Denver and leads a team of data scientists at a Denver-based media technology company.Table of ContentsPART 1 - The Building Blocks of scalable computingWhy scalable computing mattersIntroducing DaskPART 2 - Working with Structured Data using Dask DataFramesIntroducing Dask DataFramesLoading data into DataFramesCleaning and transforming DataFramesSummarizing and analyzing DataFramesVisualizing DataFrames with SeabornVisualizing location data with DatashaderPART 3 - Extending and deploying DaskWorking with Bags and ArraysMachine learning with Dask-MLScaling and deploying Dask