Practice the core questions for Python, SQL, statistics, data science, machine learning, AI engineering, LLMs and production systems. Open each question, learn the simple answer, then practice saying it naturally.
121 questions shown
A list is mutable, so you can add, remove or change items. A tuple is immutable after creation. I use lists when data must change and tuples when I want a fixed collection.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
A dictionary stores key-value pairs. It gives fast lookup by key and is useful for structured information such as user settings or model metrics.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
A set stores unique values. It is useful for removing duplicates and for fast membership checks.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
It is a compact way to build a list from another iterable, often with a transformation or condition.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
== compares values. is checks whether two variables refer to the exact same object in memory.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
*args collects extra positional arguments into a tuple. **kwargs collects extra keyword arguments into a dictionary.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
A lambda is a small anonymous function usually used for short operations such as sorting keys or simple transformations.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Exception handling uses try, except, else and finally so a program can handle errors without crashing unexpectedly.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
A generator yields values one at a time instead of storing everything in memory, which is useful for large datasets or streams.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
A shallow copy copies the outer object but can share nested objects. A deep copy recursively copies nested objects too.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
A decorator wraps a function or class to add behavior without rewriting the original implementation.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
A context manager manages setup and cleanup around a block, commonly with the with statement for files or database connections.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
PEP 8 is the common Python style guide. It improves consistency and readability across projects.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
__init__ initializes a new class instance after it is created, usually by assigning initial attributes.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Inheritance lets one class reuse and extend behavior from another class.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Polymorphism means different objects can respond to the same interface or method in their own way.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
They isolate project dependencies so package versions from one project do not break another.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
The Global Interpreter Lock in CPython allows only one thread to execute Python bytecode at a time. It affects CPU-bound threading but not all forms of concurrency.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Threads share memory and are useful for I/O-bound work. Processes have separate memory and can use multiple CPU cores for CPU-bound work.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Type hints improve readability, editor support and static checking. They do not normally enforce types at runtime.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
INNER JOIN returns only matching rows from both tables. LEFT JOIN returns every row from the left table and matching rows from the right, otherwise NULL.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
A primary key uniquely identifies each row in a table and should not be duplicated.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
A foreign key links one table to another and helps enforce referential integrity.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Normalization organizes data to reduce duplication and update problems, usually by splitting information into related tables.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Denormalization intentionally duplicates or combines data to improve read performance when the trade-off is acceptable.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
WHERE filters rows before aggregation. HAVING filters groups after GROUP BY.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
GROUP BY combines rows with the same grouping values so aggregate functions like COUNT, SUM or AVG can be calculated.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
A window function calculates across related rows without collapsing them into one row, for example ROW_NUMBER or running totals.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
A Common Table Expression is a named temporary result used inside a query, often to improve readability or support recursion.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
An index is a data structure that speeds up lookups but uses storage and can slow inserts or updates.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
A transaction groups operations that should succeed or fail together, commonly described using ACID properties.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Atomicity, Consistency, Isolation and Durability. These properties describe reliable database transactions.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
DELETE removes selected rows, TRUNCATE removes all rows efficiently, and DROP removes the table object itself.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
NULL represents missing or unknown data. It is not the same as zero or an empty string.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Group by the candidate duplicate columns and use HAVING COUNT(*) > 1.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
A subquery is a query nested inside another query. It can be used for filtering, calculation or derived tables.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
A view is a saved query exposed like a virtual table. It can simplify repeated logic or control access.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Cardinality can refer to the number of distinct values in a column or the relationship type between tables.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
A Series is one labeled dimension. A DataFrame is a two-dimensional labeled table made of columns.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
First understand why values are missing. Then choose to remove, impute, model explicitly or leave them, based on context and bias risk.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
The mean is the arithmetic average. The median is the middle value and is usually more robust to outliers.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Standard deviation measures how spread out values are around the mean.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Variance is the average squared distance from the mean. Standard deviation is the square root of variance.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Correlation means variables move together. Causation means a change in one actually produces a change in another. Correlation alone does not prove causation.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
A p-value measures how compatible observed data are with a null hypothesis under the test assumptions. It is not the probability that the null hypothesis is true.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
A confidence interval gives a range produced by a method designed to capture the true parameter at a stated long-run rate.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Sampling bias happens when the sample is not representative because some members of the target population are more or less likely to be included.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Leakage happens when training uses information that would not truly be available at prediction time, giving unrealistically strong performance.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Exploratory Data Analysis is the process of understanding distributions, relationships, missing values, anomalies and data quality before modeling.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Feature engineering creates or transforms input variables so a model can represent useful patterns more effectively.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
It converts a categorical value into separate binary indicator columns.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Standardization usually transforms a feature to have mean zero and standard deviation one.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
It often means scaling values to a fixed range or scaling each sample to a unit norm, depending on context.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
An outlier is an observation unusually far from the main distribution. It may be an error, rare event or valid extreme case.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Class imbalance means one target class has far fewer examples than another, which can make accuracy misleading.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
It counts true positives, true negatives, false positives and false negatives for a classifier.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Precision asks how many predicted positives were correct. Recall asks how many actual positives were found.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
F1 is the harmonic mean of precision and recall and is useful when balancing both matters.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
ROC-AUC measures how well a classifier ranks positive examples above negative examples across thresholds.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
A baseline is a simple reference model used to check whether a more complex model actually adds value.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Supervised learning trains on labeled targets. Unsupervised learning finds structure in data without a provided target.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Classification predicts categories. Regression predicts continuous numeric values.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Overfitting happens when a model learns training-specific noise and performs worse on unseen data.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Underfitting happens when the model is too simple or insufficiently trained to capture important patterns.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
High bias means the model is too simple. High variance means it is too sensitive to training data. Good generalization balances both.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Cross-validation repeatedly splits data into training and validation portions to estimate generalization more reliably.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Training fits the model, validation guides choices, and the final test estimates performance on untouched data.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Regularization discourages overly complex solutions to reduce overfitting, for example L1 or L2 penalties.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
L1 can push coefficients exactly to zero and encourage sparsity. L2 smoothly shrinks coefficients.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
It models the target as a weighted linear combination of features and usually chooses weights that minimize squared error.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
It models the log-odds of a class as a linear function and converts the result into a probability using the logistic function.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
A decision tree repeatedly splits data using feature rules to create groups with more similar targets.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
A random forest averages many randomized decision trees to reduce variance and improve robustness.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Gradient boosting builds weak learners sequentially, with each new learner focusing on errors from the current ensemble.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Random forest builds many trees mostly independently. XGBoost builds boosted trees sequentially and includes optimization and regularization features.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
KNN predicts using the labels or values of nearby training examples under a chosen distance metric.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
K-means assigns points to k clusters by iteratively updating cluster centers and assignments to reduce within-cluster distance.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Principal Component Analysis projects data onto orthogonal directions that capture the most variance.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
It searches settings not learned directly from training, such as tree depth or learning rate, using validation performance.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Grid search tests a predefined combination grid. Random search samples combinations and is often more efficient when only some hyperparameters matter.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Early stopping ends training when validation performance stops improving, helping reduce overfitting.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Scaling matters for algorithms based on distance or gradient optimization because features with large numeric ranges can dominate.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Feature importance estimates how much each input contributes to model predictions, but interpretation depends on the method.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Calibration checks whether predicted probabilities match observed frequencies, such as predictions near 0.8 being correct about 80% of the time.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Concept drift happens when relationships between inputs and targets change over time, reducing production performance.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
AI is the broad field of building intelligent behavior. Machine learning is a subset that learns patterns from data. Deep learning is a subset of ML using multi-layer neural networks.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
A neural network is a layered function made of weighted transformations and nonlinear activations that learns parameters from data.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Backpropagation efficiently computes gradients of the loss with respect to model parameters so an optimizer can update them.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
One epoch means the model has processed the full training dataset once.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Batch uses the whole dataset for one update. Mini-batch uses smaller groups and is the common practical approach.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
An embedding is a dense numeric vector representing semantic or structural properties so similar items can be close in vector space.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
A transformer is a neural architecture built around attention mechanisms that process relationships between tokens efficiently.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Attention lets a model weight which other tokens or features are most relevant when computing a representation.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
A large language model is a neural language model trained on large text collections to predict and generate sequences of tokens.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Tokenization converts text into smaller units called tokens that a language model can process.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
The context window is the amount of token information a model can consider in a single request.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Hallucination is when the model generates unsupported or incorrect information as if it were reliable.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Retrieval-Augmented Generation retrieves relevant external information and adds it to the model context before generation.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
A vector database stores embeddings and supports similarity search for retrieving semantically related content.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Cosine similarity compares the angle between vectors, often used to measure embedding similarity independent of magnitude.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Prompt engineering designs instructions, context and output constraints so a model has a clearer task contract.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
An AI agent combines a model with a loop for planning or deciding, using tools, observing results and continuing toward a goal.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Tool calling lets a model request a structured action, such as querying a database or API, and then use the returned result.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Fine-tuning updates model weights on task-specific examples to adapt behavior or domain performance.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Fine-tuning changes model parameters. RAG supplies external knowledge at request time. They solve different problems and can be combined.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Quantization stores or computes model weights with lower numerical precision to reduce memory and often improve inference speed.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Distillation trains a smaller student model to imitate a larger teacher model.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Inference is using a trained model to make predictions or generate outputs on new inputs.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Latency is the time from request to result. In AI systems it can include network, retrieval, model generation and post-processing time.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Throughput measures how much work a system completes per unit time, such as requests or tokens per second.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
I would combine task-specific test cases, correctness or groundedness checks, safety tests, latency, cost, user feedback and regression evaluation.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Use better instructions, trusted retrieval, constrained outputs, source grounding, verification steps and clear fallback behavior when evidence is weak.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
MLOps is the practice of reliably developing, deploying, monitoring and updating machine learning systems.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Monitor service health, latency, errors, data quality, input drift, prediction distributions, model performance where labels exist, and business outcomes.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Package preprocessing and model logic, expose an API or batch job, define dependencies, test it, deploy to infrastructure, then monitor and version it.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Docker packages an application with its runtime dependencies into containers so it behaves consistently across environments.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Continuous Integration and Continuous Delivery automate testing and deployment so changes can be released more safely and repeatedly.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
An API is a defined interface that lets software systems communicate through structured requests and responses.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
REST is request-response and fits many APIs. WebSocket keeps a persistent two-way connection for real-time communication.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Caching stores reusable results so repeated requests can be served faster and with less computation.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.
Rate limiting controls how many requests a client can make in a time window to protect reliability, cost and abuse limits.
Speaking tip: Answer in your own words, then add one short example from a project or real situation.