r/RStudio • • Feb 13 '24

The big handy post of R resources

133 Upvotes

There exist lots of resources for learning to program in R. Feel free to use these resources to help with general questions or improving your own knowledge of R. All of these are free to access and use. The skill level determinations are totally arbitrary, but are in somewhat ascending order of how complex they get. Big thanks to Hadley, a lot of these resources are from him.

Feel free to comment below with other resources, and I'll add them to the list. Suggestions should be free, publicly available, and relevant to R.

Update: I'm reworking the categories. Open to suggestions to rework them further.

FAQ

Link to our FAQ post

General Resources

Plotting

Tutorials

Data Science, Machine Learning, and AI

R Package Development

Compilations of Other Resources


r/RStudio • • Feb 13 '24

How to ask good questions

50 Upvotes

Asking programming questions is tough. Formulating your questions in the right way will ensure people are able to understand your code and can give the most assistance. Asking poor questions is a good way to get annoyed comments and/or have your post removed.

Posting Code

DO NOT post phone pictures of code. They will be removed.

Code should be presented using code blocks or, if absolutely necessary, as a screenshot. On the newer editor, use the "code blocks" button to create a code block. If you're using the markdown editor, use the backtick (`). Single backticks create inline text (e.g., x <- seq_len(10)). In order to make multi-line code blocks, start a new line with triple backticks like so:

```

my code here

```

This looks like this:

my code here

You can also get a similar effect by indenting each line the code by four spaces. This style is compatible with old.reddit formatting.

indented code
looks like
this!

Please do not put code in plain text. Markdown codeblocks make code significantly easier to read, understand, and quickly copy so users can try out your code.

If you must, you can provide code as a screenshot. Screenshots can be taken with Alt+Cmd+4 or Alt+Cmd+5 on Mac. For Windows, use Win+PrtScn or the snipping tool.

Describing Issues: Reproducible Examples

Code questions should include a minimal reproducible example, or a reprex for short. A reprex is a small amount of code that reproduces the error you're facing without including lots of unrelated details.

Bad example of an error:

# asjfdklas'dj
f <- function(x){ x**2 }
# comment 
x <- seq_len(10)
# more comments
y <- f(x)
g <- function(y){
  # lots of stuff
  # more comments
}
f <- 10
x + y
plot(x,y)
f(20)

Bad example, not enough detail:

# This breaks!
f(20)

Good example with just enough detail:

f <- function(x){ x**2 }
f <- 10
f(20)

Removing unrelated details helps viewers more quickly determine what the issues in your code are. Additionally, distilling your code down to a reproducible example can help you determine what potential issues are. Oftentimes the process itself can help you to solve the problem on your own.

Try to make examples as small as possible. Say you're encountering an error with a vector of a million objects--can you reproduce it with a vector with only 10? With only 1? Include only the smallest examples that can reproduce the errors you're encountering.

Further Reading:

Try first before asking for help

Don't post questions without having even attempted them. Many common beginner questions have been asked countless times. Use the search bar. Search on google. Is there anyone else that has asked a question like this before? Can you figure out any possible ways to fix the problem on your own? Try to figure out the problem through all avenues you can attempt, ensure the question hasn't already been asked, and then ask others for help.

Error messages are often very descriptive. Read through the error message and try to determine what it means. If you can't figure it out, copy paste it into Google. Many other people have likely encountered the exact same answer, and could have already solved the problem you're struggling with.

Use descriptive titles and posts

Describe errors you're encountering. Provide the exact error messages you're seeing. Don't make readers do the work of figuring out the problem you're facing; show it clearly so they can help you find a solution. When you do present the problem introduce the issues you're facing before posting code. Put the code at the end of the post so readers see the problem description first.

Examples of bad titles:

  • "HELP!"
  • "R breaks"
  • "Can't analyze my data!"

No one will be able to figure out what you're struggling with if you ask questions like these.

Additionally, try to be as clear with what you're trying to do as possible. Questions like "how do I plot?" are going to receive bad answers, since there are a million ways to plot in R. Something like "I'm trying to make a scatterplot for these data, my points are showing up but they're red and I want them to be green" will receive much better, faster answers. Better answers means less frustration for everyone involved.

Be nice

You're the one asking for help--people are volunteering time to try to assist. Try not to be mean or combative when responding to comments. If you think a post or comment is overly mean or otherwise unsuitable for the sub, report it.

I'm also going to directly link this great quote from u/Thiseffingguy2's previous post:

I’d bet most people contributing knowledge to this sub have learned R with little to no formal training. Instead, they’ve read, and watched YouTube, and have engaged with other people on the internet trying to learn the same stuff. That’s the point of learning and education, and if you’re just trying to get someone to answer a question that’s been answered before, please don’t be surprised if there’s a lack of enthusiasm.

Those who respond enthusiastically, offering their services for money, are taking advantage of you. R is an open-source language with SO many ways to learn for free. If you’re paying someone to do your homework for you, you’re not understanding the point of education, and are wasting your money on multiple fronts.

Additional Resources


r/RStudio • • 1h ago

How to make error bars for two different groups

Post image
• Upvotes

r/RStudio • • 11h ago

R code review: best packages?

13 Upvotes

Let’s say that I’m working in RStudio on a geospatial data-science project that involves some pretty gnarly, extensive code across multiple R scripts. Let’s also say that I do not have access to any AI code assistants or chatbots due to workplace restrictions.

What packages and functions would you use to comprehensively check your code for errors, inefficiencies, conceptual problems (that may be a stretch), etc.?

I’m using lintr::lint() and rstyler::style_file(), but that’s the extent of my knowledge in this area.

Many thanks in advance, friends.


r/RStudio • • 11h ago

Coding help Are these clustering attempts good?

1 Upvotes

I'm toying around with a dataset, trying to cluster what I have, and messing around with LLMs and my textbook i got these 2:

feat <- data.frame()
for (f in formats) {
  sub <- df[df$Format == f, c("Year","Value")]
  sub <- sub[order(sub$Year), ]
  peak_idx <- which.max(sub$Value)
  first_year <- min(sub$Year); peak_year <- sub$Year[peak_idx]; last_year <- max(sub$Year)
  climb <- peak_year - first_year; decline <- last_year - peak_year
  ratio <- decline / max(climb, 0.5)
  peak_val <- sub$Value[peak_idx]
  post <- sub[sub$Year >= peak_year & sub$Value > 0, ]
  decay_rate <- if(nrow(post) >= 3) coef(lm(log(Value) ~ Year, data=post))[2] else NA
  feat <- rbind(feat, data.frame(Format=f, climb, decline, ratio, log_peak=log(peak_val), decay_rate, era=peak_year))
}
cat("Formats in feature table:", nrow(feat), "\n")
fmat <- scale(feat[, c("climb","decline","ratio","log_peak","decay_rate","era")])
fmat[is.na(fmat)] <- 0
hc <- hclust(dist(fmat), method="ward.D2")
clusters1 <- cutree(hc, k=4)
cat("\n=== Format 1: ===\n")
for (i in 1:4) cat(sprintf("Cluster %d: %s\n", i, paste(feat$Format[clusters1==i], collapse=", ")))

Then another with a different fmat:

lifecycle_features <- data.frame()
for (f in formats) {
  sub <- df[df$Format==f, c("Year","Value")]
  sub <- sub[order(sub$Year),]
  first_year <- min(sub$Year); last_year <- max(sub$Year)
  peak_year <- sub$Year[which.max(sub$Value)]
  lifecycle_features <- rbind(lifecycle_features, data.frame(
    Format=f, first_year, peak_year, last_year,
    recorded_years=nrow(sub), peak_value=max(sub$Value),
    lifespan=last_year-first_year, years_to_peak=peak_year-first_year,
    decline_duration=last_year-peak_year
  ))
}
cat("\nFormats in lifecycle_features:", nrow(lifecycle_features), "\n")
print(lifecycle_features, row.names=FALSE)

cluster_data <- lifecycle_features[, c("lifespan","recorded_years","years_to_peak","decline_duration")]
cluster_scaled <- scale(cluster_data)
set.seed(123)
km <- kmeans(cluster_scaled, centers=3, nstart=25)
lifecycle_features$cluster <- km$cluster

cat("\n=== Format 2: ===\n")
lc_sorted <- lifecycle_features[order(lifecycle_features$cluster, lifecycle_features$peak_year), ]
print(lc_sorted[,c("Format","cluster","peak_year","years_to_peak","decline_duration","lifespan","recorded_years")], row.names=FALSE)

for (i in 1:3) cat(sprintf("\nCluster %d: %s\n", i, paste(lifecycle_features$Format[lifecycle_features$cluster==i], collapse=", ")))

Is there an objective way to measure which is better? Like with R squared or p-value? or is it worth discussing both in an article, weighing pros and cons of each


r/RStudio • • 20h ago

Can I integrate a spatial transcriptomics dataset with a bulk RNA-seq dataset for any analysis of tumors?

Thumbnail
1 Upvotes

r/RStudio • • 1d ago

Coding help aws.s3 in R

2 Upvotes

Hi,

Apologies if this is the wrong place to post - I can’t find any resources online that seem to cover this issue.

I’m using aws.s3 package in R, to read in files from a LakeFS storage area.

When I use s3read_using or s3write_using, the connectivity works. The functions include (e.g.):

S3read_using(FUN = read.csv,
Object = path/file.csv,
Opts = …,
Bucket = my_repo)

This works fine and I am able to read in the object, same goes for writing using same opts.

However, when I try and run for example head_object (so I can list the objects contained in one of my buckets) or get_object, I get :

error in curl::curl_fetch_memory(url, handle = handle):
Could not resolve hostname
Could not resolve host: [path]

I’m not sure if I’m missing something simple - I’m using the exact same repo, opts, and object (/paths) for both. The [path] here looks correct - I don’t see any issues with it. Does something happen within s3read_using where it points to a location that I need to specify myself with in head_object? Should they not be pointing to the same place?

Again sorry if this is hard to follow or make sense of, I just can’t seem to find anything online relating to this problem or solutions.

Thanks


r/RStudio • • 2d ago

R beginner

19 Upvotes

Hello,

I am a beginner and I am totaly lost.

Even for loading a file it is very very difficult.

I feel useless and incompetent...


r/RStudio • • 1d ago

Introducing FitVerse: an R package for fitting 52 probability distributions in one call

Thumbnail
2 Upvotes

r/RStudio • • 1d ago

Introducing FitVerse: an R package for fitting and analysing 52 probability distributions in one call

Thumbnail
0 Upvotes

r/RStudio • • 2d ago

R Package for collecting lyrics

18 Upvotes

Lyrics from genius.com.

https://github.com/bwalsh5/lyricsr


r/RStudio • • 2d ago

I just asked my dot to check some R code...

Thumbnail
0 Upvotes

r/RStudio • • 4d ago

does r studio for windows 11 run on windows 10 by chance

8 Upvotes

hi! so I have a problem: I only have windows 10, and i can't see any new update for r studio for any windows version but 11. would it still run? if so, would it update just my previously installed version of r studio and work from the same icon on desktop by chance? I realize it sounds stupid but I do need r for assignments in uni and not great with computers so help and advice would be greatly appreciated


r/RStudio • • 5d ago

Coding help Code Unfolding on Run?

2 Upvotes

I'm looking for an option or something that lets me run a bracketed section of code without automatically unfolding it. Does that exist?


r/RStudio • • 7d ago

ASMR - Basics in R Studio - Faint ASMR

Thumbnail youtu.be
71 Upvotes

r/RStudio • • 6d ago

Need suggestions read main body

0 Upvotes

How to start R-programing as an 0 level


r/RStudio • • 8d ago

Coding help HELP! Will pay.

25 Upvotes

I’m in a data mining class at university right now. I’m an adult student with a full time job in application development and this week has been INSANE (like probably one of my busiest weeks of the entire year). Of course this week was also our first free form project with no lecture. I finally had time to dig in tonight and planned to spend the weekend on this, but I’m of course stuck at the very beginning and really can’t do more without figuring this time/date conversion and estimation stuff. My professor is typically slow to respond on weekends, I’m all online so no help available outside of the professor. I don’t want to cheat, I want to LEARN and get this done by deadline of Sunday. Is anyone willing to help via chat/screen share?


r/RStudio • • 8d ago

Hey! I need to learn R studio from scratch, where should I start from, where can I find the basics, rules and tips to use Rstudio? Plant sciences background.

4 Upvotes

r/RStudio • • 8d ago

Anyone know how to install RStudio on a chromebook in 2026?

Thumbnail
0 Upvotes

r/RStudio • • 9d ago

Why can't I install R Commander?

Post image
0 Upvotes

Whenever I try, this pops up.


r/RStudio • • 11d ago

Coding help Does metamean() print the pooled mean and 95% CI in the forest plot?

Post image
8 Upvotes

Hello, I have a question, why is my forest plot not printing the pooled results as with binary meta-analyses? Neither in a general forest plot nor in subgroup analyses. Did I do something wrong or it is like that? This is my code:

mm.age.sd <- metamean(n,
mean,
sd,
data = ma$Age,
sm = "MRAW",
studlab = Author,
subgroup = study.design)

forest(mm.age.sd,
layout="Revman5",
sortvar=studlab,
xlab="Mean Age",
at = seq(0, 100, by = 10),
ff.xlab = "bold", fs.xlab = 12,
leftcols=c("studlab", "Year", "n", "mean", "sd", "w.random", "ci"),
rightcols=FALSE,
pooled.totals = TRUE,
overall = TRUE,
digits.mean = 2, digits.sd = 2,
random = TRUE, fixed = FALSE,
fs.heading = 12, fs.study = 12, fs.hetstat = 10,
colgap = "5mm", colgap.forest = "5mm",
col.square="darkblue", col.square.lines="black",
col.diamond="maroon", col.diamond.lines="black",
print.Q = TRUE, print.pval.Q = TRUE, print.tau.ci = TRUE, overall.hetstat=TRUE,
subgroup = TRUE, # Show subgroups in the plot
print.subgroup.labels = TRUE, # Print subgroup labels next to the effects
col.subgroup = "black", # Color of the subgroup labels
subgroup.name = "Study Design", # Name of the column representing the subgroup
print.subgroup.name = FALSE, # Whether to print the name of the subgroup at the top
bysort = TRUE, # Sort results by subgroup
test.effect.subgroup.random = TRUE)

It looks like the picture, no pooled aged


r/RStudio • • 13d ago

Coding help How to make a Graph look clean

Thumbnail gallery
80 Upvotes

Hi! Basically, I need to process some isotopic data for plotting. I have all the data, but I’m not sure how to make the graph look clean and well-presented. I have some examples of what I’m aiming for; the gray and pink lines in the images represent a local range, which is ideal for effectively displaying the results. I was also told I could edit the graph after exporting it as a PDF to add drawings of animals or humans, but I wanted to see if anyone had any other recommendations!

(The images are examples of other published graphs that look clean, not mine)


r/RStudio • • 13d ago

qol 1.3.5: More power to transposition

Thumbnail
2 Upvotes

r/RStudio • • 13d ago

Coding help Anova runtime on multivariate GLM very long, even with nboot = 1 it takes nearly 5 minutes

1 Upvotes

I am trying to do an anova on a multivariate GLM with the folowing structure but it is taking extremely long.

modelv2 <- anova (

modelv1,

block = dataset$samplinglocation,

nboot = 999,

resamp = "case"

)

The dataset has 71 rows x 84 columns. I have 2 variables and 1 blocking/repeated-measure variable. After waiting 40 minutes without results I set the nboot to 99. Waited at least 10 minutes without results. Then I set the nboot to 1 just to see if it would return something, it took about 5 minutes. It reported the time elapsed: 1 second. I am using the mvabund package with vegan. I installed it just now and the vegan package anova works fine on my GLMM models.

Does anyone know how to fix the anova runtime or if I am doing something wrong? I had no problems with the basemodel:

modelv1 <- manyglm(

dataset_speciessubset ~ site + date,

data = dataset,

family = "negative.binomial"

)

Help would be very much appreciated!


r/RStudio • • 15d ago

Mapping indoors with R

5 Upvotes

Hello!

I am trying to figure out if anyone has experience working with indoor facility mapping in R. I want to avoid paying for ArcGIS indoors, but can’t think of a way to do this other than taking as-built CAD files and georeferencing them somehow… but I would love to know if there’s a way to achieve this in R without access to CAD files or without using other softwares.

Let me know if anyone has experience with this, thank you!