
۴۴٬۰۰۰تومان
نوع فایل دانلود: EPUB
پس از خرید، یک فایل EPUB دریافت میکنید.
این فایل با Calibre، Apple Books و سایر کتابخوانهای دیجیتال مناسب است.
این کتاب مقدمهای بر استنباط علّی در پایتون است، اما بهطور کلی یک کتاب مقدماتی نیست. از این جهت مقدماتی است که تمرکزم را روی کاربرد میگذارم، نه روی اثباتها و قضیههای سختگیرانه در استنباط علّی؛ همچنین هر وقت مجبور به انتخاب باشم، توضیحی سادهتر و شهودیتر را به توضیحی کاملتر اما پیچیدهتر ترجیح میدهم. اما بهطور کلی مقدماتی نیست، چون فرض میکنم خواننده از قبل با یادگیری ماشین (ML)، آمار و برنامهنویسی در پایتون آشنایی دارد. در عین حال کتاب خیلی پیشرفته هم نیست، اما بعضی اصطلاحات را به کار میبرم که بهتر است از قبل بدانید.
یک دلار اضافه برای بازاریابی آنلاین چند خریدار جدید میآورد؟ کدام مشتریها فقط وقتی خرید میکنند که کد تخفیف بگیرند؟ چطور میشود یک راهبرد قیمتگذاری بهینه تعیین کرد؟ بهترین راه برای اینکه بفهمیم اهرمهایی که در اختیار داریم چطور روی شاخصهای کسبوکاری موردنظرمان اثر میگذارند، استفاده از استنباط علّی است.
در این کتاب، ماتیوس فاکوره، دانشمند ارشد داده در نوبانک، ظرفیت تا حد زیادی استفادهنشدهٔ استنباط علّی را برای برآورد اثرات و پیامدها توضیح میدهد. مدیران، دانشمندان داده و تحلیلگران کسبوکار با روشهای کلاسیک استنباط علّی آشنا میشوند؛ روشهایی مثل کارآزماییهای تصادفی کنترلشده (آزمونهای A/B)، رگرسیون خطی، امتیاز تمایل (propensity score)، کنترلهای مصنوعی و روش تفاوت در تفاوتها (difference-in-differences). کنار هر روش، یک کاربرد صنعتی هم آمده تا مثال ملموسی برای درک بهتر باشد.
در ادامه، فهرستی غیرجامع از چیزهایی که توصیه میکنم قبل از خواندن این کتاب بلد باشید آمده است:
- آشنایی پایه با پایتون، از جمله کتابخانههای رایج علم داده: Pandas، Numpy، Matplotlib، Scikit-Learn. پیشزمینهٔ من اقتصاد است، پس لازم نیست نگران کدنویسی خیلی پیچیده یا عجیب باشید. فقط مطمئن شوید مبانی را خوب میدانید.
- آشنایی با مفاهیم پایهٔ آمار مثل توزیعها، احتمال، آزمون فرض، رگرسیون، نویز، امید ریاضی، انحراف معیار و استقلال. اگر نیاز به مرور داشته باشید، در کتاب یک مرور آماری هم آوردهام.
- آشنایی با مفاهیم پایهٔ علم داده، مثل مدل یادگیری ماشین، اعتبارسنجی متقابل، بیشبرازش و چند مدل پرکاربرد یادگیری ماشین (گرادیان بوستینگ، درخت تصمیم، رگرسیون خطی، رگرسیون لجستیک).
با این کتاب، شما:
یاد میگیرید چطور از مفاهیم پایهٔ استنباط علّی استفاده کنید
یاد میگیرید یک مسئلهٔ کسبوکار را به مسئلهٔ استنباط علّی صورتبندی کنید
میفهمید سوگیری چگونه جلوی استنباط علّی را میگیرد
یاد میگیرید اثرات علّی چطور میتوانند از فردی به فرد دیگر متفاوت باشند
یاد میگیرید با استفاده از مشاهدههای تکراری از مشتریان یکسان در گذر زمان، سوگیریها را تعدیل کنید
میفهمید اثرات علّی در مناطق جغرافیایی مختلف چگونه تفاوت دارند
سوگیری ناشی از عدم تبعیت (noncompliance bias) و رقیقشدن اثر (effect dilution) را بررسی میکنید
مخاطب اصلی این کتاب دانشمندان دادهای هستند که در صنعت کار میکنند. اگر شما هم در این گروه هستید، احتمال زیادی دارد پیشنیازهایی را که گفتم داشته باشید. همچنین یادتان باشد این مخاطب گسترده است و مهارتهای بسیار متنوعی دارد. به همین دلیل ممکن است گاهی نکته یا پاراگرافی بیاورم که بیشتر برای خوانندهٔ پیشرفتهتر نوشته شده باشد. پس اگر همهٔ خطهای کتاب را کامل متوجه نشدید، نگران نباشید. باز هم میتوانید چیزهای زیادی از آن یاد بگیرید. و شاید بعد از اینکه بعضی از مبانیاش را بهتر یاد گرفتید، دوباره سراغش بیایید.
This book is an introduction to Causal Inference in Python, but it is not an introductory book in general. It’s introductory because I’ll focus on application, rather than rigorous proofs and theorems of causal inference; additionally, when forced to choose, I’ll opt for a simpler and intuitive explanation, rather than a complete and complex one. It is not introductory in general because I’ll assume some prior knowledge about Machine Learning (ML), statistics and programming in Python. It is not too advanced either, but I will be throwing in some terms that you should know beforehand.How many buyers will an additional dollar of online marketing bring in? Which customers will only buy when given a discount coupon? How do you establish an optimal pricing strategy? The best way to determine how the levers at our disposal affect the business metrics we want to drive is through causal inference.In this book, author Matheus Facure, senior data scientist at Nubank, explains the largely untapped potential of causal inference for estimating impacts and effects. Managers, data scientists, and business analysts will learn classical causal inference methods like randomized control trials (A/B tests), linear regression, propensity score, synthetic controls, and difference-in-differences. Each method is accompanied by an application in the industry to serve as a grounding example.Here is a non exhaustive list of the things I recommend you know before reading this book:- Basic knowledge of Python, including the most commonly used data scientists libraries: Pandas, Numpy, Matplotlib, Scikit-Learn. I come from an Economics background, so you don’t have to worry about me using very fancy code. Just make sure you know the basics pretty well.- Knowledge of basic statistical concepts like distributions, probability, hypothesis testing, regression, noise, expected values, standard deviation, independence. I will include a statistical review in the book in case you need a refresher.- Knowledge of basic data science concepts, like machine learning model, cross validation, overfitting and some of the most used machine learning models (gradient boosting, decision trees, linear regression, logistic regression).With this book, you will:Learn how to use basic concepts of causal inferenceFrame a business problem as a causal inference problemUnderstand how bias gets in the way of causal inferenceLearn how causal effects can differ from person to personUse repeated observations of the same customers across time to adjust for biasesUnderstand how causal effects differ across geographic locationsExamine noncompliance bias and effect dilutionThe main audience of this book is Data Scientists who are working in the industry. If you fit this description, there is a pretty good chance that you cover the prerequisites that I’ve mentioned. Also, keep in mind that this is a broad audience, with very diverse skill sets. For this reason, I might include some note or paragraph which is meant for the most advanced reader. So don’t worry if you don’t understand every single line in this book. You’ll still be able to extract a lot from it. And maybe it will come back for a second read once you mastered some of its basics.