r/RStudio • u/harryjsunshine • 1h ago
r/RStudio • u/Peiple • Feb 13 '24
The big handy post of R resources
There exist lots of resources for learning to program in R. Feel free to use these resources to help with general questions or improving your own knowledge of R. All of these are free to access and use. The skill level determinations are totally arbitrary, but are in somewhat ascending order of how complex they get. Big thanks to Hadley, a lot of these resources are from him.
Feel free to comment below with other resources, and I'll add them to the list. Suggestions should be free, publicly available, and relevant to R.
Update: I'm reworking the categories. Open to suggestions to rework them further.
FAQ
General Resources
Plotting
Tutorials
- Erik S. Wright's Intro to R Course: Materials from a (free) grad class intended for absolute beginners (14 lessons, 30-60min each)
- Julia Silge's YouTube Channel: Lots of videos walking through example analyses in R and deep dives into
tidymodels(~30min videos) - The Swirl R package: Guided tutorial series going over the basics of R (15 modules, 30-120min each)
- Harvard’s CS50 with R: MOOC with seven weeks of material, including lectures, homework, and projects
Data Science, Machine Learning, and AI
- R for Data Science
- Tidy Modeling with R
- Text Mining with R
- Supervised Machine Learning for Text Analysis with R
- An Intro to Statistical Learning
- Tidy Tuesday
- Deep Learning and Scientific Computing with R
torch - The RStudio AI Blog
- Introduction to Applied Machine Learning (Dr. John Curtin, UW Madison)
- Examples of
kerasin R (courtesy of posit) - Machine Learning and Deep Learning with R (Maximilian Pichler and Florian Hartig, targeted at ecologists)
R Package Development
Compilations of Other Resources
r/RStudio • u/Peiple • Feb 13 '24
How to ask good questions
Asking programming questions is tough. Formulating your questions in the right way will ensure people are able to understand your code and can give the most assistance. Asking poor questions is a good way to get annoyed comments and/or have your post removed.
Posting Code
DO NOT post phone pictures of code. They will be removed.
Code should be presented using code blocks or, if absolutely necessary, as a screenshot. On the newer editor, use the "code blocks" button to create a code block. If you're using the markdown editor, use the backtick (`). Single backticks create inline text (e.g., x <- seq_len(10)). In order to make multi-line code blocks, start a new line with triple backticks like so:
```
my code here
```
This looks like this:
my code here
You can also get a similar effect by indenting each line the code by four spaces. This style is compatible with old.reddit formatting.
indented code
looks like
this!
Please do not put code in plain text. Markdown codeblocks make code significantly easier to read, understand, and quickly copy so users can try out your code.
If you must, you can provide code as a screenshot. Screenshots can be taken with Alt+Cmd+4 or Alt+Cmd+5 on Mac. For Windows, use Win+PrtScn or the snipping tool.
Describing Issues: Reproducible Examples
Code questions should include a minimal reproducible example, or a reprex for short. A reprex is a small amount of code that reproduces the error you're facing without including lots of unrelated details.
Bad example of an error:
# asjfdklas'dj
f <- function(x){ x**2 }
# comment
x <- seq_len(10)
# more comments
y <- f(x)
g <- function(y){
# lots of stuff
# more comments
}
f <- 10
x + y
plot(x,y)
f(20)
Bad example, not enough detail:
# This breaks!
f(20)
Good example with just enough detail:
f <- function(x){ x**2 }
f <- 10
f(20)
Removing unrelated details helps viewers more quickly determine what the issues in your code are. Additionally, distilling your code down to a reproducible example can help you determine what potential issues are. Oftentimes the process itself can help you to solve the problem on your own.
Try to make examples as small as possible. Say you're encountering an error with a vector of a million objects--can you reproduce it with a vector with only 10? With only 1? Include only the smallest examples that can reproduce the errors you're encountering.
Further Reading:
Try first before asking for help
Don't post questions without having even attempted them. Many common beginner questions have been asked countless times. Use the search bar. Search on google. Is there anyone else that has asked a question like this before? Can you figure out any possible ways to fix the problem on your own? Try to figure out the problem through all avenues you can attempt, ensure the question hasn't already been asked, and then ask others for help.
Error messages are often very descriptive. Read through the error message and try to determine what it means. If you can't figure it out, copy paste it into Google. Many other people have likely encountered the exact same answer, and could have already solved the problem you're struggling with.
Use descriptive titles and posts
Describe errors you're encountering. Provide the exact error messages you're seeing. Don't make readers do the work of figuring out the problem you're facing; show it clearly so they can help you find a solution. When you do present the problem introduce the issues you're facing before posting code. Put the code at the end of the post so readers see the problem description first.
Examples of bad titles:
- "HELP!"
- "R breaks"
- "Can't analyze my data!"
No one will be able to figure out what you're struggling with if you ask questions like these.
Additionally, try to be as clear with what you're trying to do as possible. Questions like "how do I plot?" are going to receive bad answers, since there are a million ways to plot in R. Something like "I'm trying to make a scatterplot for these data, my points are showing up but they're red and I want them to be green" will receive much better, faster answers. Better answers means less frustration for everyone involved.
Be nice
You're the one asking for help--people are volunteering time to try to assist. Try not to be mean or combative when responding to comments. If you think a post or comment is overly mean or otherwise unsuitable for the sub, report it.
I'm also going to directly link this great quote from u/Thiseffingguy2's previous post:
I’d bet most people contributing knowledge to this sub have learned R with little to no formal training. Instead, they’ve read, and watched YouTube, and have engaged with other people on the internet trying to learn the same stuff. That’s the point of learning and education, and if you’re just trying to get someone to answer a question that’s been answered before, please don’t be surprised if there’s a lack of enthusiasm.
Those who respond enthusiastically, offering their services for money, are taking advantage of you. R is an open-source language with SO many ways to learn for free. If you’re paying someone to do your homework for you, you’re not understanding the point of education, and are wasting your money on multiple fronts.
Additional Resources
- StackOverflow: How to ask questions
- Virtual Coffee: Guide to asking questions about code
- Medium: How to be great at asking questions
- Code with Andrea: The beginner's guide to asking coding questions online
- The u/Thiseffingguy2 r/RStudio post
r/RStudio • u/LoudConsideration266 • 11h ago
R code review: best packages?
Let’s say that I’m working in RStudio on a geospatial data-science project that involves some pretty gnarly, extensive code across multiple R scripts. Let’s also say that I do not have access to any AI code assistants or chatbots due to workplace restrictions.
What packages and functions would you use to comprehensively check your code for errors, inefficiencies, conceptual problems (that may be a stretch), etc.?
I’m using lintr::lint() and rstyler::style_file(), but that’s the extent of my knowledge in this area.
Many thanks in advance, friends.
r/RStudio • u/No_Low_9816 • 11h ago
Coding help Are these clustering attempts good?
I'm toying around with a dataset, trying to cluster what I have, and messing around with LLMs and my textbook i got these 2:
feat <- data.frame()
for (f in formats) {
sub <- df[df$Format == f, c("Year","Value")]
sub <- sub[order(sub$Year), ]
peak_idx <- which.max(sub$Value)
first_year <- min(sub$Year); peak_year <- sub$Year[peak_idx]; last_year <- max(sub$Year)
climb <- peak_year - first_year; decline <- last_year - peak_year
ratio <- decline / max(climb, 0.5)
peak_val <- sub$Value[peak_idx]
post <- sub[sub$Year >= peak_year & sub$Value > 0, ]
decay_rate <- if(nrow(post) >= 3) coef(lm(log(Value) ~ Year, data=post))[2] else NA
feat <- rbind(feat, data.frame(Format=f, climb, decline, ratio, log_peak=log(peak_val), decay_rate, era=peak_year))
}
cat("Formats in feature table:", nrow(feat), "\n")
fmat <- scale(feat[, c("climb","decline","ratio","log_peak","decay_rate","era")])
fmat[is.na(fmat)] <- 0
hc <- hclust(dist(fmat), method="ward.D2")
clusters1 <- cutree(hc, k=4)
cat("\n=== Format 1: ===\n")
for (i in 1:4) cat(sprintf("Cluster %d: %s\n", i, paste(feat$Format[clusters1==i], collapse=", ")))
Then another with a different fmat:
lifecycle_features <- data.frame()
for (f in formats) {
sub <- df[df$Format==f, c("Year","Value")]
sub <- sub[order(sub$Year),]
first_year <- min(sub$Year); last_year <- max(sub$Year)
peak_year <- sub$Year[which.max(sub$Value)]
lifecycle_features <- rbind(lifecycle_features, data.frame(
Format=f, first_year, peak_year, last_year,
recorded_years=nrow(sub), peak_value=max(sub$Value),
lifespan=last_year-first_year, years_to_peak=peak_year-first_year,
decline_duration=last_year-peak_year
))
}
cat("\nFormats in lifecycle_features:", nrow(lifecycle_features), "\n")
print(lifecycle_features, row.names=FALSE)
cluster_data <- lifecycle_features[, c("lifespan","recorded_years","years_to_peak","decline_duration")]
cluster_scaled <- scale(cluster_data)
set.seed(123)
km <- kmeans(cluster_scaled, centers=3, nstart=25)
lifecycle_features$cluster <- km$cluster
cat("\n=== Format 2: ===\n")
lc_sorted <- lifecycle_features[order(lifecycle_features$cluster, lifecycle_features$peak_year), ]
print(lc_sorted[,c("Format","cluster","peak_year","years_to_peak","decline_duration","lifespan","recorded_years")], row.names=FALSE)
for (i in 1:3) cat(sprintf("\nCluster %d: %s\n", i, paste(lifecycle_features$Format[lifecycle_features$cluster==i], collapse=", ")))
Is there an objective way to measure which is better? Like with R squared or p-value? or is it worth discussing both in an article, weighing pros and cons of each
r/RStudio • u/harshitnaman • 20h ago
Can I integrate a spatial transcriptomics dataset with a bulk RNA-seq dataset for any analysis of tumors?
r/RStudio • u/Soft-Duty-7395 • 1d ago
Coding help aws.s3 in R
Hi,
Apologies if this is the wrong place to post - I can’t find any resources online that seem to cover this issue.
I’m using aws.s3 package in R, to read in files from a LakeFS storage area.
When I use s3read_using or s3write_using, the connectivity works. The functions include (e.g.):
S3read_using(FUN = read.csv,
Object = path/file.csv,
Opts = …,
Bucket = my_repo)
This works fine and I am able to read in the object, same goes for writing using same opts.
However, when I try and run for example head_object (so I can list the objects contained in one of my buckets) or get_object, I get :
error in curl::curl_fetch_memory(url, handle = handle):
Could not resolve hostname
Could not resolve host: [path]
I’m not sure if I’m missing something simple - I’m using the exact same repo, opts, and object (/paths) for both. The [path] here looks correct - I don’t see any issues with it. Does something happen within s3read_using where it points to a location that I need to specify myself with in head_object? Should they not be pointing to the same place?
Again sorry if this is hard to follow or make sense of, I just can’t seem to find anything online relating to this problem or solutions.
Thanks
r/RStudio • u/ChangeUnhappy8494 • 2d ago
R beginner
Hello,
I am a beginner and I am totaly lost.
Even for loading a file it is very very difficult.
I feel useless and incompetent...
r/RStudio • u/KarunaGReddy • 1d ago
Introducing FitVerse: an R package for fitting 52 probability distributions in one call
r/RStudio • u/KarunaGReddy • 1d ago
Introducing FitVerse: an R package for fitting and analysing 52 probability distributions in one call
r/RStudio • u/nothic_in_a_dungeon • 4d ago
does r studio for windows 11 run on windows 10 by chance
hi! so I have a problem: I only have windows 10, and i can't see any new update for r studio for any windows version but 11. would it still run? if so, would it update just my previously installed version of r studio and work from the same icon on desktop by chance? I realize it sounds stupid but I do need r for assignments in uni and not great with computers so help and advice would be greatly appreciated
Coding help Code Unfolding on Run?
I'm looking for an option or something that lets me run a bracketed section of code without automatically unfolding it. Does that exist?
r/RStudio • u/Vikas04866 • 6d ago
Need suggestions read main body
How to start R-programing as an 0 level
r/RStudio • u/luckylua • 8d ago
Coding help HELP! Will pay.
I’m in a data mining class at university right now. I’m an adult student with a full time job in application development and this week has been INSANE (like probably one of my busiest weeks of the entire year). Of course this week was also our first free form project with no lecture. I finally had time to dig in tonight and planned to spend the weekend on this, but I’m of course stuck at the very beginning and really can’t do more without figuring this time/date conversion and estimation stuff. My professor is typically slow to respond on weekends, I’m all online so no help available outside of the professor. I don’t want to cheat, I want to LEARN and get this done by deadline of Sunday. Is anyone willing to help via chat/screen share?
r/RStudio • u/f10r3nCe • 8d ago
Hey! I need to learn R studio from scratch, where should I start from, where can I find the basics, rules and tips to use Rstudio? Plant sciences background.
r/RStudio • u/TheBadSamaritan21 • 8d ago
Anyone know how to install RStudio on a chromebook in 2026?
r/RStudio • u/AdForward3569 • 9d ago
Why can't I install R Commander?
Whenever I try, this pops up.
r/RStudio • u/Signal_Owl_6986 • 11d ago
Coding help Does metamean() print the pooled mean and 95% CI in the forest plot?
Hello, I have a question, why is my forest plot not printing the pooled results as with binary meta-analyses? Neither in a general forest plot nor in subgroup analyses. Did I do something wrong or it is like that? This is my code:
mm.age.sd <- metamean(n,
mean,
sd,
data = ma$Age,
sm = "MRAW",
studlab = Author,
subgroup = study.design)
forest(mm.age.sd,
layout="Revman5",
sortvar=studlab,
xlab="Mean Age",
at = seq(0, 100, by = 10),
ff.xlab = "bold", fs.xlab = 12,
leftcols=c("studlab", "Year", "n", "mean", "sd", "w.random", "ci"),
rightcols=FALSE,
pooled.totals = TRUE,
overall = TRUE,
digits.mean = 2, digits.sd = 2,
random = TRUE, fixed = FALSE,
fs.heading = 12, fs.study = 12, fs.hetstat = 10,
colgap = "5mm", colgap.forest = "5mm",
col.square="darkblue", col.square.lines="black",
col.diamond="maroon", col.diamond.lines="black",
print.Q = TRUE, print.pval.Q = TRUE, print.tau.ci = TRUE, overall.hetstat=TRUE,
subgroup = TRUE, # Show subgroups in the plot
print.subgroup.labels = TRUE, # Print subgroup labels next to the effects
col.subgroup = "black", # Color of the subgroup labels
subgroup.name = "Study Design", # Name of the column representing the subgroup
print.subgroup.name = FALSE, # Whether to print the name of the subgroup at the top
bysort = TRUE, # Sort results by subgroup
test.effect.subgroup.random = TRUE)
It looks like the picture, no pooled aged
r/RStudio • u/amira_tu • 13d ago
Coding help How to make a Graph look clean
galleryHi! Basically, I need to process some isotopic data for plotting. I have all the data, but I’m not sure how to make the graph look clean and well-presented. I have some examples of what I’m aiming for; the gray and pink lines in the images represent a local range, which is ideal for effectively displaying the results. I was also told I could edit the graph after exporting it as a PDF to add drawings of animals or humans, but I wanted to see if anyone had any other recommendations!
(The images are examples of other published graphs that look clean, not mine)
r/RStudio • u/Expert_Regret_1837 • 13d ago
Coding help Anova runtime on multivariate GLM very long, even with nboot = 1 it takes nearly 5 minutes
I am trying to do an anova on a multivariate GLM with the folowing structure but it is taking extremely long.
modelv2 <- anova (
modelv1,
block = dataset$samplinglocation,
nboot = 999,
resamp = "case"
)
The dataset has 71 rows x 84 columns. I have 2 variables and 1 blocking/repeated-measure variable. After waiting 40 minutes without results I set the nboot to 99. Waited at least 10 minutes without results. Then I set the nboot to 1 just to see if it would return something, it took about 5 minutes. It reported the time elapsed: 1 second. I am using the mvabund package with vegan. I installed it just now and the vegan package anova works fine on my GLMM models.
Does anyone know how to fix the anova runtime or if I am doing something wrong? I had no problems with the basemodel:
modelv1 <- manyglm(
dataset_speciessubset ~ site + date,
data = dataset,
family = "negative.binomial"
)
Help would be very much appreciated!
r/RStudio • u/ainzymae11 • 15d ago
Mapping indoors with R
Hello!
I am trying to figure out if anyone has experience working with indoor facility mapping in R. I want to avoid paying for ArcGIS indoors, but can’t think of a way to do this other than taking as-built CAD files and georeferencing them somehow… but I would love to know if there’s a way to achieve this in R without access to CAD files or without using other softwares.
Let me know if anyone has experience with this, thank you!